YC Summer 2026 (S26) company

Conifer

Least cost routing system to reduce 70%+ token spend

About Conifer

There are thousands of models, providers, and harnesses, each with its own pricing and strengths, leaving an overwhelming number of choices when it comes to managing token spend and intelligence. Conifer uses intelligent routing and orchestration logic to handle the full path of each query: model selection, provider choice, and cache management. It all runs inside your existing harness, so tools like Claude Code and Codex work without changing your setup. By centralizing where inference occurs, Conifer lets companies save on their inference bill while leaving teams free to focus on what they are building.

Public traction evidence

Each signal links to the public source used for attribution.

  1. X

    Conifer's SDK is open-sourced as of today.

    Conifer's SDK is open-sourced as of today. It is one gateway for every kind of inference you would need: cloud or self-hosted models, fine-tuned models, your own API keys, as well as the hardware that is under your desk. Every model works through this after a simple install in

  2. GitHub

    ConiferKit/sage

    ConiferKit/sage. YC S26 snapshot lists Conifer with official GitHub org https://github.com/ConiferKit; repo lives under that org.

  3. X

    Little peak into what we've been working on: - Run any model from any harness from one API key (claude code, codex, open code, pi, etc) - Access to new fusion models pushing the cost and performance frontier - Local-cloud routing with configurable model pools! Very very soon

    Little peak into what we've been working on: - Run any model from any harness from one API key (claude code, codex, open code, pi, etc) - Access to new fusion models pushing the cost and performance frontier - Local-cloud routing with configurable model pools! Very very soon

  4. X

    Conifer routes 80 percent of AI requests to local hardware

    Conifer routes 80 percent of AI requests to local hardware

  5. X

    Local has been the trend for this week.

    Local has been the trend for this week. JetBrains launched Junie Local on the 24th, with Perplexity following soon with their own Portable Computer launched on the 25th. Both of these run the entire agent on the hardware using a 27B Qwen model. The interesting part of these

  6. X

    The team at @TryOpenTag is truly N of 1 It's great to see more apps building around model agnostic endpoints and we're happy we get to be a part of their journey.

    The team at @TryOpenTag is truly N of 1 It's great to see more apps building around model agnostic endpoints and we're happy we get to be a part of their journey. Check them out at https://t.co/TrE6LDoNwx and if you're looking to hedge against singleton model buildout check out

  7. X

    Today we’re bringing a better user experience to model selection and agent runtime.

    Today we’re bringing a better user experience to model selection and agent runtime. Everyone agrees you should use the best model. But the current infrastructure makes this impossible. Look at what actually exists: the leading harnesses offer a handful of models. Open-weight and

  8. X

    Reading over an old report from @EpochAIResearch.

    Reading over an old report from @EpochAIResearch. The trend lines suggest we'll achieve a model of comparable intelligence to Fable/Sol that you'll be able to host on a single consumer GPU.