YC Summer 2026 (S26) company

Robocurve

Evals for robots

About Robocurve

Robocurve builds open-source tools and independent benchmarks to measure how well robots can do real-world jobs. Instead of relying on unverified demo videos from frontier labs, we score their models on reproducible benchmarks that anyone can trust. Today there are no well-run, standardized robotics benchmarks. Labs evaluate in-house, and no independent group has stepped in to run continuous benchmarking as a service. The result is that no one actually knows how good anyone else is, or where the real frontier sits. Good benchmarks require operating and maintaining physical hardware and real-world setups. Simulation only goes so far, since models that look strong in sim can show large performance gaps once deployed in the real world. Building real-world benchmarks means coordinating job-domain experts, evals engineering, and hands-on robotics all at once. Frontier labs are targeting general-purpose robotics by 2028, yet the field of robotics evals barely exists. Whoever builds the trusted measure of robot capability becomes the reference everyone relies on. We combine backgrounds in AI evals and robotics to build and grow this field as fast as possible. We already shipped v1 of our open-source framework, Inspect Robots, and ran our first pilot scoring a frontier model on a real robot.

Public traction evidence

Each signal links to the public source used for attribution.

  1. X

    We gave Opus 5 robot arms.

    We gave Opus 5 robot arms. It can stack bowls.

  2. X

    Gemini 3.7 Flash just saturated one of our physical tool-use benchmarks at 92%.

    Gemini 3.7 Flash just saturated one of our physical tool-use benchmarks at 92%. Gemini 3.6 Flash, released three weeks earlier, scored only 32%. This marks a step change in the robotics capabilities of LLMs. 🧵

  3. X

    We evaluated DeepMind's new Gemini Robotics-ER 2 model.

    We evaluated DeepMind's new Gemini Robotics-ER 2 model. It threw our robot arms off the table. Here's what happened 🧵

  4. X

    Introducing @robocurve, a Public Benefit Corporation to measure and report frontier robotics capabilities.

    Introducing @robocurve, a Public Benefit Corporation to measure and report frontier robotics capabilities. We build real-world evaluations for robots and publish results as a neutral third party.

  5. X

    Opus 5 can do this with zero demonstrations despite not being trained for robots.

    Opus 5 can do this with zero demonstrations despite not being trained for robots. On the flip side Opus took 7 minutes while GEN-1.5 took 7 seconds. 🧵

  6. X

    Robots can bake bread, cut carrots, and do karate kicks in demo videos. But how capable are they, really? @robocurve measures how good A...

    Robots can bake bread, cut carrots, and do karate kicks in demo videos. But how capable are they, really? @robocurve measures how good AI and robots are in the physical world, from making a sandwich to building data centers.

  7. X

    Jay - a Rhodes Scholar from Harvard - is launching today with a new take on how we should measure progress in robotics. Congrats on the l…

    Jay - a Rhodes Scholar from Harvard - is launching today with a new take on how we should measure progress in robotics. Congrats on the launch @chooi_jeq !

  8. X

    Yes, YC does back Public Benefit Corporations https://t.co/bzrvWF1jnJ

    Yes, YC does back Public Benefit Corporations https://t.co/bzrvWF1jnJ