Each Monday, the evaluation pass compares current test outputs, usage costs, and acceptance criteria across the Eden AI models under consideration. An unreviewed model, routing, or budget change can raise spend sharply or put materially different outputs into a production workflow.
01
A model choice needs a testable reason.
Strawberry can collect the provider, model, task, price, latency, and output evidence visible in the Eden AI and project tabs you open. It turns a scattered comparison into a reviewable decision brief.
02
Usage numbers become useful when they are attached to a job.
A dashboard total does not reveal which workflow created the spend or whether the output justified it. Strawberry can prepare a usage review tied to the use cases and tests in your browser context.
03
Test results should be readable by the person making the call.
Strawberry can turn a batch of visible outputs and evaluation notes into a concise summary of quality, failure modes, and next experiments without overstating what the tests prove.
04
Evaluations lose value when every run starts from zero.
Each Monday, the recurring pass assembles the latest test outputs, usage notes, and acceptance criteria into a model-comparison packet. Switching a model, changing routing, or raising a budget without review can alter production output quality and inflate usage costs.
It can prepare model comparisons, usage reviews, test summaries, and evaluation packets from the signed-in context you choose.
No. Provider, configuration, spending, and production decisions require explicit approval.
Yes. It can organise visible tests and notes into a reviewable summary, with limits and open questions made clear.
Yes. Each Monday, it can rebuild the model scorecard from current tests, usage, and acceptance criteria while leaving model, routing, and budget changes for review.