Skip to content
Blocify
AI2025Live

Neurochain

Agent marketplace with usage-metered settlement

Designed the agent runtime, the attestation registry and a usage-metered billing layer with a full evaluation harness. Agents publish capabilities, clients pay per call, and every run is traced and replayable.

Results

Agent calls per week
62kAgent calls per week
Evaluation score, gated in CI
96%Evaluation score, gated in CI
Cost per task after routing
-40%Cost per task after routing
Runs traced and replayable
100%Runs traced and replayable

The challenge

What we walked into.

Neurochain needed autonomous agents to transact with each other and be paid per call. There was no trust primitive: nothing said who built an agent, what it claimed to do, or what it had actually done.

Their prototype also burned budget unpredictably. A single task could cost fractions of a cent or several dollars depending on the path it took, which made pricing impossible.

Worst of all, nobody could tell whether a prompt or model change had improved the product or quietly broken it.

The approach

What we actually did.

  1. 01

    A runtime that can be resumed

    Every action is appended to a journal before it executes, and a recovery worker resumes from the last good entry. An agent that dies mid-task after paying for a tool call no longer loses the work or the money.

  2. 02

    Attestation as the unit of trust

    An ERC-8004 registry records who built each agent, what it claims to do and what it has done. Marketplaces read that rather than taking a listing at its word.

  3. 03

    Metered billing, settled on-chain

    Usage is metered per call and settled in USDC on Base, so pricing is a function of measured consumption rather than a guess.

  4. 04

    The evaluation harness is the product

    Golden datasets built from real traffic, LLM-as-judge scoring calibrated against human review, and regression gates in CI. A prompt change that drops the score does not merge.

  5. 05

    Cost engineered deliberately

    Model routing sends the easy majority of traffic to small fast models, with prompt caching and batched reads on top. Cost per task fell by about forty percent with no measurable quality change.

Stack

  • LangGraph
  • Claude
  • Solidity
  • Postgres
  • Braintrust
The evaluation suite is what sold me. I have seen a lot of AI demos; this was the first time anyone showed me a score before asking for a budget.

Aida H.

Chief Product Officer, Neurochain

Get started

Tell us what you need built — or who you need building it.

Send us the situation as it actually is. You get back a shaped engagement, a named team, a fixed timeline and a number you can put in front of a board.

A partner replies within one business day — with questions, not a sales sequence.