When the Judge Acts: Auditing VLM-Guided Image Selection on Culturally Situated Prompts

Huichan SeoIndependent Researcher

Same images, different order

A vision-language judge is asked to return the image that best fits the request Chilean family lunch, children eagerly finishing homemade pastel de choclo. We show it the same three images in three orders.

SD-3.5-Large image for the Chilean family lunch prompt; human prompt-alignment rating 0.00.
SD-3.5rated 0.00
Flux.1-Dev image for the Chilean family lunch prompt; human prompt-alignment rating 0.50.
Flux.1rated 0.50
GPT-Image image for the Chilean family lunch prompt showing pastel de choclo; human prompt-alignment rating 0.83.
GPT-Imagerated 0.83
0.00

Human rating of the image the judge returns

Qwen3-VL-4B takes slot A in every order, so the order alone decides whether the user gets an image rated 0.00 or 0.83; random choice averages 0.44. Across all 300 prompts, reordering changes the returned image on 59.7%.

Three findings

  1. +0.039

    Barely better than random

    Gain in human rating (0–1 scale) when Qwen3-VL-4B picks, over random choice from the same images. A CLIP similarity baseline, blind to order, reaches +0.069.

    Results
  2. 48.8%

    Strong position bias

    of Qwen3-VL-4B’s 900 calls pick the first image shown, against 27.8% if it chose uniformly. At 8B the share falls to 27.3%, close to uniform.

    Position bias
  3. +0.141

    Agreement helps if it targets the failure

    gain for the 4B on the 40.3% of prompts where all three orders agree; a weaker second judge instead keeps below-random picks (−0.060). At 8B, the prompts this filter rejects still beat random (+0.068).

    Agreement filters

Abstract

Vision-language models (VLMs) are increasingly used as judges that pick the best of several generated images, so their choices decide what users see. Such judges are usually validated by how well their scores agree with human ratings, not by the images they choose. We audit VLM judges as decision-makers: on 300 culturally situated prompts, we compare the image a judge returns with human ratings it never sees and with random choice from the same candidates, and we repeat every decision with the candidates reordered. A 4B-parameter judge barely beats random selection and falls short of a CLIP similarity baseline. Its choices show strong position bias: it picks the first image shown in 49% of calls, where chance would give 28%, and reordering changes the image it returns on 60% of prompts. For this judge, agreement across orders is informative: decisions that survive reordering are much better than random, whereas agreement with a weaker second judge keeps the wrong ones. An 8B judge from the same family shows almost no position bias and beats the CLIP baseline, and for it the same order-agreement filter mostly discards good decisions. Agreement helps only when it targets the judge’s failure mode, so such filters must be re-audited whenever the judge changes. The 4B judge’s slight rise in stereotype ratings is no longer detectable after aggregating across orders or with the larger judge.

How the audit works

The judge never sees a human rating. We record what it picks, and only afterwards compare that pick with the ratings and with random choice from the same images. Scroll to follow one prompt through the audit.

Step 1 of 6
  1. Step 1 of 6

    300 prompts, each with 3–4 generated images

    Each square is one prompt about everyday culture, 30 per country. Every prompt comes with 3 or 4 images from different generators, fixed before any judging.

  2. Step 2 of 6

    Show the same images in three orders

    The judge gets the prompt and the images, nothing else. We ask three times and rotate the images, so each one takes a turn in slot A.

  3. Step 3 of 6

    The judge picks one image per order

    Qwen3-VL-4B picks slot A all three times, so it returns a different image each time. SmolVLM2-2.2B picks B, B, then A.

  4. Step 4 of 6

    Agreement filter: act or abstain

    The filter acts only if all three orders return the same image. Here they don’t, so it abstains. Across all prompts it keeps 121 of 300.

  5. Step 5 of 6

    Only now, reveal the human ratings

    Human ratings stayed hidden while the judge chose. We attach them by image ID only after every pick is recorded.

  6. Step 6 of 6

    Score the pick against random and the best

    Regret is the gap between the pick and the best image; random is the average of the pool. Here the order alone moves the result from 0.00 to 0.83.

See the full protocol diagram (Figure 1 in the paper)

Position bias: where the judge looks

Share of choices by presentation slot. Dashes mark the rate expected if the judge chose uniformly, given the mix of three- and four-image pools. The ablations change the prompt format and keep the bias.

Qwen3-VL-4B (900 calls) favours slot A; SmolVLM2-2.2B (886 calls) favours slots A and B and rarely picks later slots. Both depart sharply from uniform (χ²3 = 219 and 522).

Does scale remove the bias?

After the main audit we re-ran all 900 calls with Qwen3-VL-8B from the same family, and with an 8-bit copy of the 4B as a control for quantization and runtime. Each column is one judge, on the same 300 prompts and the same three orders.

Top: share of calls choosing the first slot; dashes mark uniform choice. Bottom: gain over random choice in human rating with 95% bootstrap intervals, for every prompt (original order), for prompts where all three orders return the same image, and for the rest; the dashed line is the CLIP baseline. The 8-bit 4B control returns the same image as the main 4B run on 94.9% of calls, so the 8B differences are unlikely to come from quantization or the runtime. Post hoc; 8B and control run on Apple MLX with 8-bit language-model weights.

At 8B the position bias nearly disappears, and the order-agreement filter mostly discards good decisions: the prompts it rejects still beat random by +0.068. Agreement helps only when it targets the judge’s failure mode.

Agreement filters: what they keep and reject

An agreement filter lets the judge act only on prompts where its choices agree, and abstain on the rest. Each dot is one of the 300 prompts, one row per country. The first three filters use Qwen3-VL-4B; the last applies the order filter to the 8B judge. Judge a filter by both sides: kept prompts should beat random, and rejected ones should not.

Kept 121 of 300 prompts (40.3%)

kept rejected

Gain over random choice in human rating (95% CI)

Kept (121)
+0.141
Rejected (179)
−0.030

The order filter keeps a prompt only if Qwen returns the same image in all three orders. Kept prompts beat random by +0.141; the prompts it rejects fall below random (−0.030).

Results

Table 1. Gain over random choice on the same images, in human prompt-alignment rating (95% CI).
PolicyPromptsGain
CLIP (content only)300+0.069 [0.047, 0.090]
Qwen, original order300+0.039 [0.016, 0.062]
Qwen, order majority300+0.065 [0.043, 0.086]
Qwen, filter: 3/3 orders agree121+0.141 [0.110, 0.173]
rejected by that filter179−0.030 [−0.057, −0.002]
Qwen, filter: both judges agree59−0.060 [−0.111, −0.011]
rejected by that filter241+0.064 [0.039, 0.087]
Smol, original order297−0.073 [−0.093, −0.052]

“Rejected” rows apply Qwen’s original-order choice to the prompts a filter declines. These rows and CLIP are exploratory.

Paired differences between Qwen's choices and random choice in three cultural-error ratings: missing explicit expectations decreases, missing implicit expectations changes little, stereotype rating rises slightly.
Figure 2. Cultural-error ratings of Qwen’s choices minus random choice; negative is fewer errors. Qwen reduces missing explicit expectations (−0.056) but slightly raises stereotype ratings (+0.023, pointwise CI [0.002, 0.045]); the Bonferroni-adjusted interval [−0.002, 0.051] includes zero, so this is exploratory.

The 4B’s small stereotype rise does not survive aggregation. Taking the majority choice across the three orders gives +0.007 [−0.013, 0.027]; keeping only unanimous prompts gives +0.010 [−0.019, 0.039]; the 8B judge’s original-order choices give +0.016 [−0.004, 0.036]. All three intervals include zero.

Selection regret by country group for random, Qwen and Smol across ten countries, with bootstrap intervals.
Figure 3. Country-group regret with 95% bootstrap intervals (30 prompts per group). Random regret shows each group’s room for improvement. Country groups are dataset contexts, not homogeneous cultural preferences.

Examples

Each gallery is chosen by a fixed rule over the recorded outputs, not by looking at the images. Numbers are mean human prompt-alignment ratings (0–1).

Four CulturalFrames prompts from Iran, Poland, China and Brazil. In each, Qwen's pick, framed in red, is rated below the pool mean while a clearly better candidate, framed in green, is available.
Figure 4. VLM judges can return below-average images when a clearly better candidate exists. Red: Qwen3-VL-4B’s original-order choice, below the pool mean. Green: highest-rated candidate. The examples illustrate the failure mode, not its frequency.
Five prompts, one per row, where Qwen picks slot A in all three orders and so returns different images; the last column shows the best-rated candidate.
Figure 5. Order flips. Prompts where Qwen picks slot A in all three orders and returns at least two different images; the five with the widest rating range among its picks. The last column is the best-rated candidate.
Prompts where CLIP's choice and Qwen's choice differ most in human rating: on the left CLIP's choice is rated higher, on the right Qwen's is.
Figure 6. Content scorer versus judge. Prompts where CLIP’s choice (C) and Qwen’s original-order choice (Q) differ most in rating: left, CLIP’s is rated higher; right, Qwen’s is. Checks mark the best-rated candidates. Rows with three images are prompts that have only three candidates.
Prompts where Qwen's choice has the highest stereotype rate in its pool although another candidate with equal or higher alignment has a lower stereotype rate.
Figure 7. Stereotype trade-off. Qwen’s choice has the highest stereotype rate in its pool, and another candidate with equal or higher alignment has a lower rate (checks). Labels give mean alignment (A) and the share of annotators flagging a stereotype (S). Rows with three images are prompts that have only three candidates.
Unfiltered examples: the first four-candidate prompt for each of ten countries, candidates in original order, with Qwen and Smol choices framed and the highest-rated candidates checked.
Figure 8. Unfiltered examples: the first four-candidate prompt per country, candidates in original order. Blue (Q) and orange (S) frames mark the Qwen and Smol original-order choices; green checks mark the highest-rated candidates.

Audit browser

The released artifact includes a static browser over every recorded decision: headline estimates, per-prompt pools with both judges’ choices in each order, and a filterable prompt explorer, all computed from the recorded outputs.

Screenshot of the audit browser overview page showing headline estimates, slot preference, country-group regret and agreement-filter charts.
Figure 9. Overview: headline estimates, presentation-slot preference, country-group regret, kept-versus-rejected gains for each agreement filter, and gain over random for every policy.
Screenshot of the audit browser detail view for the Chilean family lunch prompt, showing three orders with each judge's choice and per-candidate ratings.
Figure 10. Detail view for the demo prompt: the three orders with each judge’s choice and per-candidate human ratings.
Screenshot of the audit browser prompt explorer filtered to Qwen's below-mean choices and sorted by regret.
Figure 11. Prompt explorer filtered to Qwen’s below-mean choices, sorted by regret.

Limitations and data use

The main audit covers two small open judges, 300 prompts, and one set of released human ratings. Country groups are dataset contexts, not homogeneous cultural preferences. Position, complement-set, CLIP, ablation and stereotype comparisons, and the 8B scale extension (run with 8-bit weights on a different runtime, with an 8-bit 4B control), were added after main inference and are exploratory. All images and ratings come from CulturalFrames (Nayak et al., Findings of EMNLP 2025); images appear here only to illustrate the analysis. Please obtain the data from its original release.

BibTeX

@misc{seo2026judgeacts,
  title  = {When the Judge Acts: Auditing VLM-Guided Image Selection
            on Culturally Situated Prompts},
  author = {Seo, Huichan},
  year   = {2026},
  note   = {Preprint},
  url    = {https://seochan99.github.io/JudgeActs/}
}