YC Spring 2026 (P26) company

Miso Labs

Emotive foundation voice models

b2binfrastructuredeveloper tools7 public signals

About Miso Labs

Miso Labs is building the world’s most emotive foundation models for voice. We believe that the next generation of AI interactions shouldn't just be functional—they should be human. By bringing warmth and lightning-fast speed to the voice layer, we empower developers to build voice agents that users truly love.

Public traction evidence

Each signal links to the public source used for attribution.

  1. GitHub

    MisoLabsAI/MisoTTS

    MisoLabsAI/MisoTTS. YC S2026 snapshot lists Miso Labs with X handle MisoLabsAI; repo owner MisoLabsAI and repo title identify Miso TTS.

  2. YouTube

    Launching Miso One: The Most Emotive AI Voice Model

    Today, we’re excited to introduce Miso One, the most emotive voice model in the world. Miso One is an 8-billion-parameter text-to-speech model for highly expressive speech generation. It emotes like a human and responds faster than a human, with just 110 milliseconds of...

  3. YouTube

    Miso One Demo: TikTok Voiceover

    This entire voiceover was generated by Miso One, the most emotive voice model in the world. Miso One is an 8-billion-parameter text-to-speech model for highly expressive speech generation. It emotes like a human and responds faster than a human, with just 110 milliseconds of...

  4. YouTube

    Miso One Demo: Sports Voiceover

    This entire voiceover was generated by Miso One, the most emotive voice model in the world. Miso One is an 8-billion-parameter text-to-speech model for highly expressive speech generation. It emotes like a human and responds faster than a human, with just 110 milliseconds of...

  5. X

    “AI contributed basically zero to GDP” says less about AI than about GDP.

    “AI contributed basically zero to GDP” says less about AI than about GDP. Even before the AI boom, GDP failed to count data, the main input to AI. This paper estimates measured GDP would be 4% to 6% higher if data were counted. Imagine how this graph would look in 2026. From

  6. YouTube

    We Cloned Sal Khan

    This entire voiceover was generated by Miso One, the most emotive voice model in the world. Miso One is an 8-billion-parameter text-to-speech model for highly expressive speech generation. It emotes like a human and responds faster than a human, with just 110 milliseconds of...

  7. X

    getting your kid a summer job is a luxury imo

    getting your kid a summer job is a luxury imo