YC Summer 2026 (S26) company

Lamb Labs

Lightning fast chips with hardcoded AI models

About Lamb Labs

Lamb Labs is building MPUs (Model Processing Units), custom chips that hardcode LLM weights directly into the chip. Unlike GPUs, which move weights between memory and compute, MPUs keep the model on-chip, targeting up to 20,000+ tok/s and a 63× higher intelligence per watt.

Public traction evidence

Each signal links to the public source used for attribution.

  1. X

    12,000 tok/s on a 250$ FPGA board.

    12,000 tok/s on a 250$ FPGA board. It's a small model, and doesn't answer coherently to any question or prompt really. try it out on https://t.co/RP2Qo6kZ3L but don't expect a good chat. at least its VERY fast

  2. X

    Lamb Labs (YC S26) is building custom AI inference chips that can do 20,000+ tok/s at 63x higher Intelligence per Watt (IPW) than traditional GPUs.

    Lamb Labs (YC S26) is building custom AI inference chips that can do 20,000+ tok/s at 63x higher Intelligence per Watt (IPW) than traditional GPUs. GPUs became the default for AI because they were available, not because they were the right hardware.

  3. X

    Meet Woolly 🐑 We post-trained Qwen3-8B to run up to 2–3× faster on math & coding prompts and, more importantly, to fit on our chip.

    Meet Woolly 🐑 We post-trained Qwen3-8B to run up to 2–3× faster on math & coding prompts and, more importantly, to fit on our chip. The technique works with any LLM, regardless of size or quantization. Live demo this week → https://t.co/7K5HUfVq5Z

  4. X

    custom chips at 20,000 + tok/s!!! building the future of hardware for ai inference 🐑🐑

    custom chips at 20,000 + tok/s!!! building the future of hardware for ai inference 🐑🐑

  5. X

    spot the lambs outside of the HQ, on a road trip to meet customers 🐑🐑

    spot the lambs outside of the HQ, on a road trip to meet customers 🐑🐑

  6. X

    Researcher ➡️ Venture Scientist Look forward to starting my journey at @conceptionxtech in Cohort 6.

    Researcher ➡️ Venture Scientist Look forward to starting my journey at @conceptionxtech in Cohort 6.

  7. LinkedIn

    Small language models are getting good enough that the next bottleneck is no longer just model quality.

    Small language models are getting good enough that the next bottleneck is no longer just model quality. It is where inference runs. A few years ago, a 7B-class model was barely useful for many real workloads. Today, that same size class is crossing into much more capable...

  8. X

    busy lambs @ycombinator

    busy lambs @ycombinator