LLLoki's Lab
Lab Notes·August 30, 2026

Fleet Skill Matrix v2 Methodology

-

DOC 1 — Fleet Skill Matrix v2 Methodology (LL-010)

1.0 Overview

This document defines the standardized procedure for generating scores within the Loki’s Lab Fleet Skill Matrix. The objective is to ensure that every benchmark result represents a comparable, reproducible measure of model capability against hardware constraints and specific task definitions. All published scores must adhere to this framework to maintain integrity and transparency.

2.0 Test Suite Definition

The v2 Methodology utilizes exactly 19 starting tests selected for stability and coverage. These tasks are categorized into logical clusters. A score is only generated if the model attempts the task successfully or fails consistently.

2.1 Logical Reasoning (3 Tests)

2.2 Coding and Technical Implementation (4 Tests)

2.3 Mathematical Computation (3 Tests)

2.4 Knowledge and RAG (4 Tests)

2.5 Language and Creativity (4 Tests)

2.6 Utility and Safety (1 Test)

3.0 Scoring Protocol

To mitigate variance caused by temperature sampling or stochastic generation, we utilize a strict median-of-three scoring protocol.

3.1 Execution Cycle

For each test ID:

  1. Run the task three times using the same seed configuration unless the task is inherently deterministic (where one run is used).
  2. Record the raw score for each execution.
  3. Sort the three scores in ascending order.
  4. Identify the median value as the official result.

3.2 Outlier Handling

If the distribution of results deviates significantly from a tight cluster, we investigate environment logs. If two runs are valid and one is a clear outlier (defined as greater than 15% deviation from the mean of the three), we recalculate using only the remaining two runs if applicable to the test definition, or default to the single run median for non-deterministic tasks.

3.3 Normalization

Raw scores are normalized against a baseline dataset published alongside the matrix release notes.

4.0 Budget Classification

All results are tagged with a hardware budget tier based on total system acquisition cost at time of build.

These tiers are used to stratify leaderboard segments and ensure fair comparison between hardware constraints.

5.0 Data Field Definitions

Every benchmark entry must populate the following fields:

5.1 Model Fields

5.2 Hardware Fields

5.3 Budget Tier

Derived from the Hardware Fields using the ranges in Section 4.0.

6.0 Coverage and N/A Handling

Coverage indicates the percentage of the 19 starting tests successfully executed on a given run.

7.0 Known Limitations

Users must acknowledge specific limitations inherent to this methodology:

8.0 Reproducibility Requirements

To challenge a published score, an independent party must be able to:

  1. Access the specific prompt set associated with the Test ID.
  2. Replicate the hardware configuration or equivalent class within the same budget tier.
  3. Run the median-of-three protocol documented in Section 3.0.
  4. Match the normalized score within a margin of error of 5% to claim statistical parity.

DOC 2 — Weekly Editorial Operating Rhythm (LL-039)

1.0 Objective

This document defines the operational cadence for Loki’s Lab content creation and moderation. The process is designed for a solo operator environment where automation handles data gathering, but human oversight ensures quality control. The goal is consistent output without sacrificing editorial integrity.

2.0 Human-in-the-Loop Policy

Automation tools are permitted to draft text, collect raw results, and summarize discussion threads. However, the human operator retains full authority over:

3.0 Weekly Cadence Breakdown

The solo operator follows this weekly rhythm to balance workload and output quality.

3.1 Monday: Trusted Feed Review

3.2 Tuesday: One Test and Configuration

3.3 Wednesday: One Article or Lesson Draft

3.4 Thursday: Submission Moderation

3.5 Friday: Newsletter and Discord Summary

3.6 Saturday: Maintenance and Cleanup

4.0 Automation Roles

To support the human operator, automation tools have restricted permissions:

The human must read and approve every draft generated by automation before it is presented to the final publication stage. This prevents hallucinated claims or biased summarization from slipping into the public record.

Jack Blair — author photo

Jack Blair

Writer & tester

Jack Blair is an independent documentary filmmaker, storyteller, and lifelong technology obsessive. Through Happy Jack Media, he explores overlooked human stories and experiments with new ways to create and connect. He founded Loki's Lab as a community where curious people can test local AI models, share what they learn, and discover what today's technology can do on the computers they already own.

View author profile →