News

How AmolfiBench is designed

A model can write an answer. Can it finish the work? AmolfiBench is being built to compare models on documents, outreach, spreadsheets, visual work, and computer use inside the same Amolfi workflow. No results are published yet.

A benchmark for work should start where work ends: with the thing a person asked for. A polished answer in chat is not the same as a finished deck, a correct spreadsheet, a useful visual, or a project updated in the right place.

AmolfiBench is how we will compare models in Amolfi’s working environment. It has not run yet, and no result from it is published. The task set is designed to span documents and decks, outreach, analysis, visual work, and computer use. Each task has a brief, a deliverable, and a way to check the result.

01 / THE UNIT OF WORK

A task ends with an artifact.

Choose a kind of work. The brief, expected output, and review change together.

THE REQUEST

Turn a client brief and a folder of source material into a six-slide pitch.

THE DELIVERABLE

A presentation someone can actually send, with claims grounded in the source files.

WHAT GETS CHECKED

Story, factual accuracy, visual hierarchy, and whether revisions follow the brief.

Illustrative task briefs show the evaluation design, not scored model runs.

One harness, across the kinds of work people do

Every candidate will get the same request, the same starting records, the same tools, and the same boundaries. A model can search, draft, revise, and use the computer available to that task. It cannot turn a required human approval into an automatic action just to finish faster.

02 / THE COMPARISON

Hold the work steady. Change the model.

01Same briefTask, source material, and starting state
02Same workspaceAmolfi tools, access, and approval boundary
03Same reviewDeliverable, action trace, and outcome

That makes the comparison about how a model performs the work, rather than which model received an easier environment.

Quality is more than a plausible answer

We will review the finished artifact and the path taken to produce it. Some checks are mechanical: a formula can be recomputed, a record change can be verified, and a missing file is simply missing. Creative work also needs human review against the brief. A model should not earn full credit for an elegant response that leaves the actual task undone.

01Completion

Did the requested deliverable exist at the end?

02Correctness

Do the facts, formulas, and actions hold up?

03Usefulness

Could a person take the result forward without starting over?

04Judgment

Did the agent stop, ask, or seek approval at the right moment?

Read the score alongside the cost of getting there

The preview chart pairs CursorBench’s coding-task score with the cost per task Cursor reports for its own coding harness. That is Cursor’s number, not an Amolfi Crown debit. Switch its horizontal axis to tokens or steps to see how much work a run used. The table keeps the full model list available, including older variants; selecting any row adds or removes its model family from the graph. When Amolfi work evaluations are ready, the same view can show their results.

A single number cannot tell you whether a model is the best choice for every job. A visually strong result, a correct financial sheet, and careful computer use ask for different strengths. The task mix, scoring rubric, model settings, sample count, and important failure modes belong beside any published score.

Explore the model comparison

Follow the work as it takes shape. Product updates, previews, and the thinking behind what we build.