01 · Dataset

Use difficult, lawful images

Each species needs multiple photographers and aquariums, with juvenile and adult animals, common colour morphs, neutral and blue lighting, partial occlusion and ordinary phone-camera blur. Add lookalike pairs, non-fish animals, species outside the catalogue and images too poor to identify. Keep one image from only one split to prevent near-duplicate leakage.

02 · Ground truth

Labels must be independent of the model

Record the accepted scientific name, life stage when known, provenance and licence. Ambiguous images should be removed or explicitly labelled unresolvable. A future public score should state who labelled the set and how disagreements were settled.

03 · Measures

Accuracy alone hides unsafe behaviour

Top-1 catalogue accuracy

Exact reviewed profile returned first.

Safe no-match rate

Unknown or unusable images rejected instead of forced into the catalogue.

False high confidence

Wrong answers labelled high confidence; the most important failure measure.

Latency and cost

Median and 95th-percentile response time plus API cost per accepted image.

04 · Release gate

Publish the failures, too

Run the frozen set without manual retries. Report every metric, model identifier, prompt version, catalogue version and test date. Retain a confusion table and representative failure categories. Do not compare models on different image sets.

Download benchmark checklist (CSV)