Co-Annotator: Expert-Distilled ViT and VLM for Visual and Documentation Guidance in Age-Related Macular Degeneration
Ziheng "Leo" Li, Benjamin Freeman, Akshay Raman, Kavin Aravindhan Rajkumar, Xinxin Fang, Rishabh Srivastava, Steven Feiner, Kaveri A. Thakoor
Deploy both visual and documentation guidance simultaneously—the combination unlocks efficiency gains neither achieves alone. Best for training environments where residents already achieve acceptable accuracy but need throughput.
Ophthalmology residents waste time searching retinal scans for pathology and documenting biomarkers. Expert gaze and dictation patterns could be distilled into AI guidance, but would it actually speed diagnosis without sacrificing accuracy?
Method: Two expert-distilled models—a gaze-aligned Vision Transformer producing fixation-aligned areas of interest and an ontology-bounded VLM pre-filling biomarker summaries—were deployed with residents. When combined, correct diagnoses per minute increased by 40% and comment editing time fell by 67%, with no accuracy loss. The striking result: neither modality improved efficiency alone during guidance (US2), but together they produced efficiency gains substantially exceeding either modality (US3).
Caveats: Tested only on retinal OCT with ophthalmology residents. Generalization to other imaging modalities or specialties unverified.
Reflections: Why did combined guidance unlock efficiency gains that neither modality achieved independently during guidance? · Do the perceptual efficiency gains from AOI guidance persist beyond the immediate post-guidance carryover period? · Would attending physicians show similar efficiency gains, or is the benefit specific to trainees?