How well do models understand a port?

PortMind evaluates vision models on port camera images. We ask models to identify trucks and container attachments, then compare their answers with human labels.

Example labels
Hover or tap an object
Port of Montréal · Viterra camera

Earlier studies

ModelAgreement with human labels

No score available: Moondream 3.1.

July 28, 2026 · 23 scored task rows · One human reference

Always answering “No” scores 65.2% here. Compare misses and false alarms below.

Study setup and limitations

20 unique images plus 4 repeat tasks. One undecidable task excluded, leaving 23 binary rows and 19 unique decidable images. Container labels include 8 positive and 15 negative rows. Repeats are not independent evidence.

Read about this study →

Misses and false alarms

In guided inspection, Llama 3.2 Vision found all 8 positive tasks but also flagged 14 of 15 negative tasks.

Guided inspection · Container trucks
0%0%50%50%100%100%Container tasks found ↑False alarms on negative images →Better ↖Grok 4.5Mistral Small 3.1Llama 4 ScoutLlama 3.2 VisionLLaVA 1.5

Select a model to open its results. 8 positive and 15 negative task rows, including repeats.

Benchmark v1.0

We are preparing a new image set with independently reviewed labels. Each model will be evaluated on the same images and scoring rules.

Montréal · Truck and container recognition · v1.0In preparation

The image set needs independent human review. Results will be published after the labels are finalized and model runs are complete.

Tasks

Truck presence

Can you see at least one truck? Pickups and larger road trucks count.

Container attachment

Is any truck carrying a shipping container or towing an empty chassis? A container standing nearby does not count.

Each question has three answers: Yes, No or Unsure. The benchmark tests what is visible in an image, not traffic flow or waiting time.

See labeling examples →

How we evaluate models

01

Label the images

Reviewers independently answer the same questions, with model answers hidden.

02

Review the answers

Reviewers resolve disagreements before we finalize the reference labels.

03

Run the models

Each model receives the same images and questions. We record its answers and any failed requests.

04

Report the results

We compare model answers with the reference labels and report agreement, misses and false alarms.

Read the method →

Contact

For questions about the benchmark or evaluating a model, get in touch.