Run the V3 Benchmark
Loki's Lab tests agents on a long-context task (V3: 4096 input tokens, 256 max output). Run it on your hardware, get reproducible results, and submit evidence to our community leaderboard.
Step 1: Choose Your Platform
Select your operating system to begin.
About the V3 Test
- Task: Long-context reasoning (4096 prompt tokens, 256 max output tokens)
- Metric: Can the model find a hidden word buried in context? (pass/fail) + quality/accuracy score (0โ5)
- Runtime: Wall-clock time in seconds (how fast the model can process the task)
- Repeatability: The test is deterministic โ same input always produces the same output
Requirements by Platform
- macOS: M1+ Mac with 16GB+ RAM, or Intel Mac (slow, CPU-only). Ollama required.
- Linux: 16GB+ system RAM. NVIDIA/AMD GPU recommended (automatic detection). Ollama required.
- Windows: Windows 10+ with native execution. GPU VRAM detection included. Ollama required.
Need Help?
Read the complete how-to guide โ