Undergraduate Research Assistant · March 2026 — present
Dr. John Virostko’s group · Dell Medical School, UT Austin
Merlin is a vision-language model: it takes a CT volume and a clinical text report, and its published results cover phenotype classification, five-year risk prediction and report generation.
I built an edge-case suite across six categories, scored on cosine similarity, logit drift, KL divergence, top-3 rank retention, entropy and confidence gap. The first thing it turned up was not a degradation curve. It was a flat line.
Fifteen text variants against one fixed CT volume — correct, wrong, empty, negated, self-contradictory, a stock-market discussion, Mandarin, typos. The image–text similarity score moves exactly as you would hope. The predictions do not move at all.
Text impact · 15 variants, image held fixedabdominal CT
| Text input | Sim | Drift | KL | Rank |
| Correct | 0.3882 | 0.0000 | 0.0000 | 3/3 |
| Wrong diagnosis | −0.0705 | 0.0000 | 0.0000 | 3/3 |
| Empty string | 0.2856 | 0.0000 | 0.0000 | 3/3 |
| Negation | −0.2256 | 0.0000 | 0.0000 | 3/3 |
| Adversarial | −0.0539 | 0.0000 | 0.0000 | 3/3 |
| Finance text | 0.0985 | 0.0000 | 0.0000 | 3/3 |
| Mandarin | 0.1294 | 0.0000 | 0.0000 | 3/3 |
Drift = mean absolute change in prediction logits against baseline. KL = shift in the full prediction distribution. Rank = top-3 conditions still matching baseline. Zero across every variant.