Carry the context.
Repeated instructions. Shared documents. Growing conversations. We’re exploring cache reuse to make returning to useful context less expensive.
CONTEXT REUSEUseful agents do more than answer once.
We’re building open-model inference for long contexts, repeated turns, and the work in between.
More useful work from every GPU.
One prompt is just the beginning. We look at the whole session—and optimize the system around it.
Repeated instructions. Shared documents. Growing conversations. We’re exploring cache reuse to make returning to useful context less expensive.
CONTEXT REUSEA fast first token matters. So does everything after it. We evaluate response time, generation speed, and behavior when requests arrive together.
SESSION PERFORMANCEThe model, engine, and hardware should work together. We measure configurations against real workload requirements, then improve what matters.
WORKLOAD-LED OPTIMIZATIONAgents return to the same instructions and context. An exact shared prefix can be reused, leaving less input to process again.
Try a few turns. See what stays useful.
The first request processes its instructions, context, and new input.
We’re interested in teams whose agents do sustained, context-heavy work. These are the workloads shaping our research.
Keep repository context in play through debugging, edits, and the next test run.
Work through substantial source material across extraction, comparison, and synthesis.
Explore and refine over repeated steps, with performance measured across the whole run.
Efficiency only counts when it holds up on the workload you actually need to run.
Context length, session patterns, concurrency, and what a good response time means for your users.
A reproducible configuration and a clear picture of quality, latency, throughput, and resource use.
Targeted changes, measured against the same work. Keep the gains that survive the comparison.
We’re building Odasher and looking for the workloads that will shape it. Start with your model, your context, and where the system slows down.
Create a brief for a future pilot. No account needed.Odasher is in development. We’re validating our serving approach before opening a production service. The pilot brief helps define a workload and its success criteria; it does not create an account or reserve GPU capacity.
Our research focuses on open-weight language models, including the GLM and DeepSeek families. The model and configuration for a future pilot will depend on workload fit, licensing, and deployment support. There is no live model catalog yet.
Time to first token, generation speed, end-to-end latency, output quality, cache reuse, and resource cost under representative load. We want results that reflect a complete session.
Not yet. We’ll share measured results with their workload and configuration once our experiments are complete. Pilot scope and commercial terms will be discussed individually.