Foresight Lab · Research

Research

005Published
FL-2026-0057 Oct 2026Report

Same model, two harnesses: ARC-AGI-3 at 4.6× fewer model calls

Agents, Benchmarks
  • 37/37 levels, both harnesses
  • 82,664 vs 103,779 output tokens
  • 184 vs 853 model calls
  • $2.42 vs $140.61 at list price
46×fewer model calls
FL-2026-0041 Oct 2026Report

16 agents build the Colosseum in Minecraft: team against solo

Agents, Multi-agent
  • 8 agents on a message board, 8 solo
  • worst team agent 0.72 vs best solo agent 0.58
  • mean gain +0.046 (p = 0.13)
  • Qwen 3.8 27B on RTX 5090s
16agents, team vs solo
FL-2026-00329 Sep 2026Note

The Open Loop: 790 self-edits in a running Lisp image

Agents, Systems
  • 16,790 commits
  • 790 self-edits, 737 in 30 days
790self-edits
FL-2026-00223 Sep 2026Report

DJev against Astra: 30 minutes of RuneScape fishing

Agents, Benchmarks
  • +1.84M XP vs +1.61M XP
  • 26B diffusion model on one RTX 5090, every action its own
  • about $0.32 of power vs $13.62 of API
  • strong on repetitive skills, weaker on abstract tasks
184MXP in 30 minutes, Astra 1.61M
FL-2026-00119 Sep 2026Note

A think budget for DJev: 79.6 to 86.9 with the same weights

Models, Benchmarks
  • 512-token think budget, was 0
  • 79.6 to 86.9 on TypeSafe's published evals
  • no fine-tuning, same base weights
869from 79.6, same weights