FL-2026-0057 Oct 2026ReportSame model, two harnesses: ARC-AGI-3 at 4.6× fewer model callsAgents, Benchmarks37/37 levels, both harnesses82,664 vs 103,779 output tokens184 vs 853 model calls$2.42 vs $140.61 at list pricearXiv ↗46×fewer model calls
FL-2026-0041 Oct 2026Report16 agents build the Colosseum in Minecraft: team against soloAgents, Multi-agent8 agents on a message board, 8 soloworst team agent 0.72 vs best solo agent 0.58mean gain +0.046 (p = 0.13)Qwen 3.8 27B on RTX 5090sarXiv ↗Post ↗16agents, team vs solo
FL-2026-00329 Sep 2026NoteThe Open Loop: 790 self-edits in a running Lisp imageAgents, Systems16,790 commits790 self-edits, 737 in 30 days790self-edits
FL-2026-00223 Sep 2026ReportDJev against Astra: 30 minutes of RuneScape fishingAgents, Benchmarks+1.84M XP vs +1.61M XP26B diffusion model on one RTX 5090, every action its ownabout $0.32 of power vs $13.62 of APIstrong on repetitive skills, weaker on abstract tasksarXiv ↗Post ↗184MXP in 30 minutes, Astra 1.61M
FL-2026-00119 Sep 2026NoteA think budget for DJev: 79.6 to 86.9 with the same weightsModels, Benchmarks512-token think budget, was 079.6 to 86.9 on TypeSafe's published evalsno fine-tuning, same base weightsarXiv ↗Post ↗869from 79.6, same weights