OPEN MODELS. FORWARD MOTION.

Inference for
agents that
keep going.

Useful agents do more than answer once.
We’re building open-model inference for long contexts, repeated turns, and the work in between.

In development Open to design partners
FIG. 01 THE INFERENCE LOOP
Sculptural green loop formed from parallel paths with a bright token flowing through it
Keep the context.
Move the work forward.
CONTEXT → COMPUTE → CONTINUATION
THE WAY WE THINK ABOUT INFERENCE

More useful work from every GPU.

01 / THE APPROACH

A better fit for
the way agents work.

One prompt is just the beginning. We look at the whole session—and optimize the system around it.

01

Carry the context.

Repeated instructions. Shared documents. Growing conversations. We’re exploring cache reuse to make returning to useful context less expensive.

CONTEXT REUSE
02

Keep a steady rhythm.

A fast first token matters. So does everything after it. We evaluate response time, generation speed, and behavior when requests arrive together.

SESSION PERFORMANCE
03

Tune the whole system.

The model, engine, and hardware should work together. We measure configurations against real workload requirements, then improve what matters.

WORKLOAD-LED OPTIMIZATION
02 / LESS REPEATING. MORE DOING.

The next turn
shouldn’t start
from zero.

Agents return to the same instructions and context. An exact shared prefix can be reused, leaving less input to process again.

Try a few turns. See what stays useful.

Reusable prefix New input
agent_sessionCONCEPTUAL EXAMPLE
USER

Find the cause of this failing test.

System instructionsNEW
Repository contextNEW
Current requestNEW
A fresh starting point.

The first request processes its instructions, context, and new input.

An illustration of prefix reuse, not a performance benchmark. Actual reuse depends on the model, serving engine, and exact input prefix.
03 / THE WORK AHEAD

Built around the work.
Not just the request.

We’re interested in teams whose agents do sustained, context-heavy work. These are the workloads shaping our research.

Coding agents

Keep repository context in play through debugging, edits, and the next test run.

MANY TURNS. ONE CODEBASE.

Document workflows

Work through substantial source material across extraction, comparison, and synthesis.

MORE CONTEXT. LESS REWORK.

Agent systems

Explore and refine over repeated steps, with performance measured across the whole run.

EVERY STEP COUNTS.
A MEASURED APPROACH

First, understand.
Then, optimize.

Efficiency only counts when it holds up on the workload you actually need to run.

  1. 01

    Map the workload

    Context length, session patterns, concurrency, and what a good response time means for your users.

  2. 02

    Establish a baseline

    A reproducible configuration and a clear picture of quality, latency, throughput, and resource use.

  3. 03

    Test the improvements

    Targeted changes, measured against the same work. Keep the gains that survive the comparison.

EARLY DAYS. AMBITIOUS WORK.

Let’s move
your work forward.

We’re building Odasher and looking for the workloads that will shape it. Start with your model, your context, and where the system slows down.

Create a brief for a future pilot. No account needed.
OPEN MODELS. PRACTICAL PROGRESS.IN DEVELOPMENT / 2026
A FEW DETAILS

Before the first turn.

Can I use Odasher today?

Odasher is in development. We’re validating our serving approach before opening a production service. The pilot brief helps define a workload and its success criteria; it does not create an account or reserve GPU capacity.

Which models are you focused on?

Our research focuses on open-weight language models, including the GLM and DeepSeek families. The model and configuration for a future pilot will depend on workload fit, licensing, and deployment support. There is no live model catalog yet.

What will you measure?

Time to first token, generation speed, end-to-end latency, output quality, cache reuse, and resource cost under representative load. We want results that reflect a complete session.

Are performance results or prices available?

Not yet. We’ll share measured results with their workload and configuration once our experiments are complete. Pilot scope and commercial terms will be discussed individually.

LET’S START WITH THE WORKLOAD

Plan your pilot.

Turn an idea into a useful starting point. This creates a brief on your device; it doesn’t submit an application.

A SMALL, SIMPLE WEBSITE

Your visit. Your data.

This website does not use analytics, advertising trackers, or third-party fonts. The pilot planner runs in your browser. Your entries are not sent to Odasher or stored by the site, and disappear when you reload or close the page.

If you download or copy your workload brief, you control where it goes next. The hosting provider may process ordinary request information, such as IP address and browser details, to deliver and secure this site.

Last updated September 8, 2026.