The Spice Lab

The intelligence behind Wasabe, harnessed in-house.

The Spice Lab is our research group. It doesn’t build models; it builds reliability. Wasabe runs on a leading American-made frontier model, post-trained and fine-tuned through Wasabe Spice Forge, our harness, until one agent can carry an assignment across tens of thousands of interactions.

10k+interactions / assignment100+tools orchestrated1agent, on course
Spice ForgeThe Spice Lab · agent harness
in production
RoutesOne right path among hundreds.
Tools100+ tools, called correctly.
VerificationA second agent checks the work.
10k+interactions / assignment100+tools1agent, on course
The research thesis

A capable model isn’t a reliable agent.

Give a brilliant model one question and it shines. Give it a real assignment (a hundred steps, a dozen apps, tools that bite back), and every step is a chance to drift. Small errors compound; at step ten thousand, “usually right” isn’t good enough.

That is the problem the Spice Lab exists to solve. Our research isn’t about making the model smarter; the raw intelligence comes from a leading American-made model. It’s about making the agent dependable: choosing among many possible routes, choosing among many possible tools, and staying on course across tens of thousands of interactions.

  • The right route, even when there are hundreds.
  • The right tool, even when there are a hundred-plus.
  • Still on course at interaction ten thousand.
Wasabe Spice Forge

Where raw capability becomes finished work.

Wasabe Spice Forge is our proprietary harness: the post-training, fine-tuning, and runtime strategy that turns a frontier model into a dependable coworker. It orchestrates research, tool discovery, tool use, and reasoning, and holds the agent steady across tens of thousands of interactions. It’s Wasabe’s own IP, and where most of the value is created.

  1. 1
    Research

    Searches the web and your workspace for what the task actually needs.

  2. 2
    Discover tools

    Picks the right tools and apps for the job from everything available.

  3. 3
    Use tools

    Calls them in order: drafting, querying, and editing across your apps.

  4. 4
    Reason

    Works over the results, plans the next move, and keeps the goal in view.

  5. 5
    Verify

    For high-stakes work, a second agent independently challenges and checks it.

  6. 6
    Deliver

    Hands back a finished, sourced result you can act on.

Where the value is createdThe model gives capability; the Forge turns it into finished work.Verification built inFor high-stakes work, a second agent challenges and checks before delivery.
The research program

Reliability, studied from every angle.

The model supplies the raw intelligence. Everything the Spice Lab studies is what it takes to make that intelligence dependable inside real work.

Route reliability

At every step of an assignment there are many possible paths. Our research is making the agent take the right one, and recover fast when it doesn’t.

Tool reliability

More than a hundred tools, each with sharp edges. We train the agent to discover, choose, and call the right one, with the right arguments, every time.

Tens of thousands of interactions

Real assignments aren’t one prompt. The Forge keeps a single agent coherent across tens of thousands of tool calls, messages, and decisions.

Post-trained for the harness

Spice Forge is a training strategy as much as a runtime. We post-train and fine-tune against the harness’s own traces, so model and harness fit hand in glove.

A leading American-made model

The raw intelligence comes from a leading American-made frontier model: world-class capability, hosted and controlled in North America.

Independent verification

For high-stakes work, a second agent challenges and checks the result before it’s delivered.

The platform is free. The agent is why you’re here.

Every Wasabe agent already runs inside Spice Forge. Claim your seat and hand it your worst Tuesday.