01 · Dataset
Use difficult, lawful images
Each species needs multiple photographers and aquariums, with juvenile and adult animals, common colour morphs, neutral and blue lighting, partial occlusion and ordinary phone-camera blur. Add lookalike pairs, non-fish animals, species outside the catalogue and images too poor to identify. Keep one image from only one split to prevent near-duplicate leakage.
02 · Ground truth
Labels must be independent of the model
Record the accepted scientific name, life stage when known, provenance and licence. Ambiguous images should be removed or explicitly labelled unresolvable. A future public score should state who labelled the set and how disagreements were settled.
03 · Measures
Accuracy alone hides unsafe behaviour
Top-1 catalogue accuracy
Exact reviewed profile returned first.
Safe no-match rate
Unknown or unusable images rejected instead of forced into the catalogue.
False high confidence
Wrong answers labelled high confidence; the most important failure measure.
Latency and cost
Median and 95th-percentile response time plus API cost per accepted image.
04 · Release gate
Publish the failures, too
Run the frozen set without manual retries. Report every metric, model identifier, prompt version, catalogue version and test date. Retain a confusion table and representative failure categories. Do not compare models on different image sets.
Download benchmark checklist (CSV)