Olam Labs
Building multi-agent simulations for model evals and training.
About Olam Labs
Multi-agent environments let us evaluate and train models in complex, simulated worlds. We work with researchers on both evaluations and training for character and agentic performance typically difficult to assess for using current popular datasets. Play one of our first releases, Multi-Agent Arena (https://olamlabs.ai/arena), where humans come play social strategy games against multiple AI agents. It's currently top 50 on OpenRouter and has thousands of matches played each day. We use Multi-Agent Arena to build datasets on agentic performance and real-world socialization.
Public traction evidence
Each signal links to the public source used for attribution.
- X
opus 4.6 is the last anthropic model that i had no problem reading the writing of quickly all of the models since, incl fable, are a lot less legible.
opus 4.6 is the last anthropic model that i had no problem reading the writing of quickly all of the models since, incl fable, are a lot less legible. for both coding and convo is it an RL artifact? what's causing it? i've seen this sentiment shared by friends
- X
NEW: Ox Alpha, the @OpenRouter stealth model, ranks #4 on our Elo Rating, just behind GPT-5.6 Sol.
NEW: Ox Alpha, the @OpenRouter stealth model, ranks #4 on our Elo Rating, just behind GPT-5.6 Sol. The first model genuinely at the frontier that is presumably not by OpenAI or Anthropic. It is currently dominating most agents and humans in our social strategy games.
- X
We're releasing Social Arena, and our first benchmark, the Deception Index Social Arena is the first platform where humans come play social games like Risk, Catan, or Poker versus AI agents These multi-agent matches are then used for evaluations on model behavior More below!
We're releasing Social Arena, and our first benchmark, the Deception Index Social Arena is the first platform where humans come play social games like Risk, Catan, or Poker versus AI agents These multi-agent matches are then used for evaluations on model behavior More below!
- X
Can you beat frontier AIs at social strategy games? Today, we're releasing Multi-Agent Arena for free public access! Play games like Risk, Catan, or Poker against GPT, Claude, etc.
Can you beat frontier AIs at social strategy games? Today, we're releasing Multi-Agent Arena for free public access! Play games like Risk, Catan, or Poker against GPT, Claude, etc. We're using this run evaluations on human and model behavior in complex environments.
- X
i believe we need much better benchmarks on how models behave for ai acceleration to go well when i was little i spent all day playing games like age of empires and minecraft so it’s a bit full circle that now that i get to spend all day building similar games as multi agent https://t.co/7Ku2HGaV7D
i believe we need much better benchmarks on how models behave for ai acceleration to go well when i was little i spent all day playing games like age of empires and minecraft so it’s a bit full circle that now that i get to spend all day building similar games as multi agent...
- X
ox alpha (the new @OpenRouter stealth model) is genuinely insane normally when new models proclaiming to be frontier come out they're spiky in capability and on our evals don't match across multi-agent arena it appears that's now changed!
ox alpha (the new @OpenRouter stealth model) is genuinely insane normally when new models proclaiming to be frontier come out they're spiky in capability and on our evals don't match across multi-agent arena it appears that's now changed!
- X
YAYY :D me + agents found a new theorem fully solving erdős problem 872, verified in lean! took ~400 prompts and ~600 commits over 3 months opus had been managing a team of 5.4 pros since april, past week of fable + 5.6 pro got a full solution https://t.co/d4kgKZEJro
YAYY :D me + agents found a new theorem fully solving erdős problem 872, verified in lean! took ~400 prompts and ~600 commits over 3 months opus had been managing a team of 5.4 pros since april, past week of fable + 5.6 pro got a full solution https://t.co/d4kgKZEJro
- X
hello i've been going around talking about multi agent envs irl and i would like to clarify the different definitions i've seen subagent orchestration (the one you know): eg claude code or codex subagents.
hello i've been going around talking about multi agent envs irl and i would like to clarify the different definitions i've seen subagent orchestration (the one you know): eg claude code or codex subagents. post training tasks where the model is incentivized to orchestrate.