UnMaskFork
Test-time scaling where multiple masked diffusion LMs collaboratively unmask an answer, guided by MCTS.
-
Interactive Research Blogs
Test-time scaling where multiple masked diffusion LMs collaboratively unmask an answer, guided by MCTS.
We replace Picbreeder's human users with AI (large Vision-Language Models), and ask which ingredients let AI search open-endedly the way people do.
A benchmark for evaluating how much net income an LLM agent can generate as a coffee roaster over 90 days in a multi-agent economy with two farmers, two roasters, and two retailers.
Watch digital species compete, cooperate, and evolve in real time.
Hypernetworks that generate LoRA adapters from documents or text descriptions.
LLMs evolve assembly programs that battle for control of a virtual machine.

Extending LLM context windows by dropping positional embeddings after pretraining.
A lightweight module that assigns semantically-informed positions to tokens, enabling better attention on noisy contexts and long sequences.
Multiple neural cellular automata species compete and adapt in a shared environment.
Why Sudoku variants remain a grand challenge in AI reasoning.
Neural networks that use timing and synchronisation as a representation for thought.
A benchmark testing creative reasoning through Sudoku variants with invented rules.
A small Japanese language model that runs entirely in the browser.
Foundation models automatically discover diverse artificial life simulations.