Oct 2, 2026 · science · 15min read

FunBench: Benchmarking Fundus Reading Skills of MLLMs — Paper Reading Notes

TL;DR

FunBench is a novel hierarchical dataset and benchmark for evaluating the effectiveness of open-source multimodal large language models (MLLMs) on fundus reading tasks of varying difficulty. The code and data are available on GitHub and Hugging Face. 1

Overview of FunBench

1. Why This Paper Matters

The paper provides a reasonable benchmark for evaluating the fundus reading abilities of MLLMs. It can be used directly to assess the performance of MLLMs or agent workflows in ophthalmology. It also provides a method and an approach for developing MLLM benchmarks in other domains. 1

2. Background and Motivation

2.1 Context and Problem

Although research on medical MLLMs is growing rapidly, the development of ophthalmology-focused benchmarks is lagging behind. 1

2.2 Gap in Prior Work

Prior medical benchmarks, such as OmniMedVQA and GMAI-MMBench, include retinal images within broader medical evaluations. However, they do not provide the particular ophthalmology-specific task hierarchy and combination of targeted module evaluations proposed by FunBench. GMAI-MMBench itself includes multiple tasks and levels of perceptual granularity, so it should not be characterized as an unstructured, single-task benchmark. 1 4 5

In its comparison with LMOD, FunBench emphasizes the recognition of major anatomical structures in fundus images and two diseases. This is not a complete account of LMOD, which also covers multimodal information and subgroup analysis. 1 6

3. Core Contributions

  • It designs hierarchical tasks ranging from low-level modality and anatomy perception to high-level lesion analysis and disease diagnosis.
  • It provides targeted proxy evaluations of two key MLLM modules—the vision encoder (VE) and the large language model (LLM)—together with a holistic evaluation. These modes offer complementary views but do not completely isolate the contributions of the modules. 1

4. Method

4.1 Overview

FunBench organizes fundus reading into four levels and evaluates MLLMs through three targeted evaluation modes. 1

4.2 Dataset Curation

4.2.1 Hierarchical Task Organization

Level 1 (L1): Modality perception

  • L1a: Coarse-grained modality perception

    Please choose the most suitable option based on the image and the question. Answer directly with the option letter.
    Question: What kind of image is presented here?
    Options: A. Fundus image; B. Remote sensing; C. Painting; D. Whole-slide image.

  • L1b: Fine-grained modality perception

    Please choose the most suitable option based on the image and the question. Answer directly with the option letter.
    Question: Which imaging technique was used to capture the image?
    Options: A. Ultra-widefield fundus photography; B. Ultrasound; C. Fluorescein fundus angiography; D. X-ray.

Level 2 (L2): Anatomy perception

  • L2a: Determine the relative position of the optic disc and fovea

    Please choose the most suitable option based on the image and the question. Answer directly with the option letter.
    Question: Which is on the left, the optic disc or the fovea?
    Options: A. Optic disc; B. Fovea.

  • L2b: Recognize the laterality of the fundus image

    Please choose the most suitable option based on the image and the question. Answer directly with the option letter.
    Question: Which eye is shown in this image, the left or the right?
    Options: A. Left eye; B. Right eye.

Level 3 (L3): Lesion analysis

  • L3a: Lesion recognition

    Please choose the most suitable option based on the image and the question. Answer directly with the option letter.
    Question: Does this fundus image show any evidence of fibrous proliferation?
    Options: A. No; B. Yes.

  • L3b: Lesion localization

    Please choose all suitable options based on the image and the question. Answer directly with the option letters, separated by commas if necessary.
    Question: Where is the fibrous proliferation in this fundus image?
    Options: A. Temporal to the optic disc center; B. Nasal to the optic disc center; C. Not observed in the image; D. Inferior to the optic disc center; E. Superior to the optic disc center.

  • L3c: Lesion size estimation

    Please choose the most suitable option based on the image and the question. Answer directly with the option letter.
    Question: What is the size of the fibrous proliferation depicted in the fundus photograph?
    Options: A. Not observed in the image; B. No more than 0.5 disc area; C. Between 0.5 and 1.0 disc area; D. Between 1.0 and 1.5 disc areas; E. Between 1.5 and 2.0 disc areas; F. More than 2.0 disc areas.

  • L3d: Lesion counting

    Please choose the most suitable option based on the image and the question. Answer directly with the option letter.
    Question: How many haemorrhages do you see in the image?
    Options: A. Not observed in the image; B. No more than 5; C. Between 5 and 15; D. Between 15 and 25; E. Between 25 and 35; F. More than 35.

Level 4 (L4): Disease diagnosis

  • L4a: Binary-condition diagnosis

    Please choose the most suitable option based on the image and the question. Answer directly with the option letter.
    Question: Does this fundus image show any signs of illness?
    Options: A. Yes; B. No.

  • L4b: Multi-condition diagnosis

    Please choose all suitable options based on the image and the question. Answer directly with the option letters, separated by commas if necessary.
    Question: What specific abnormalities are shown in this fundus image?
    Options: A. No abnormality; B. Retinitis pigmentosa; C. Retinal vein occlusion; D. Retinal artery occlusion; E. Diabetic retinopathy.

  • L4c: Diabetic retinopathy grading

    Please choose the most suitable option based on the image and the question. Answer directly with the option letter.
    Question: What stage of diabetic retinopathy can be observed in the given fundus image?
    Options: A. Proliferative diabetic retinopathy; B. Mild nonproliferative diabetic retinopathy; C. Severe nonproliferative diabetic retinopathy; D. No diabetic retinopathy; E. Moderate nonproliferative diabetic retinopathy.

  • L4d: Age-related macular degeneration categorization

    Please choose the most suitable option based on the image and the question. Answer directly with the option letter.
    Question: What is the age-related macular degeneration category of the presented fundus image?
    Options: A. Early age-related macular degeneration; B. Intermediate age-related macular degeneration; C. No age-related macular degeneration; D. Late age-related macular degeneration.

The exact released task definitions are available in the FunBench dataset. 3

4.2.2 Data Sources

FunBench adapts the following 14 public datasets: 1

  • Six CFP datasets: IDRiD, DDR, JSIEC, RFMiD, OIA-ODIR, and Retinal-Lesions
  • Five OCT datasets: OCTDL, NEH, OCTID, UCSD, and RETOUCH
  • One UWF dataset: TOP
  • Two multimodal datasets: MMC-AMD (CFP + OCT) and DeepDRiD (CFP + UWF)

Table 1. FunBench statistics: 16,348 fundus images and 91,810 visual questions across the reported task hierarchy. The rows and task identifiers enumerate 2 + 2 + 4 + 4 = 12 tasks, although the original Table 1 caption reports 10 tasks. 1 (Table 1)

LevelSingle-answer questionsMultiple-answer questionsSample questionData sources
L1
Tasks: 2
32,6960What method was used to capture this image?
A. Magnetic resonance imaging
B. Ultra-widefield fundus photography
C. Color fundus photography
D. Optical coherence tomography
All datasets
L2
Tasks: 2
10,9800Which eye is shown in this image, the left or the right?
A. Right eye
B. Left eye
CFP: DDR, DeepDRiD, IDRiD, OIA-ODIR, Retinal-Lesions
UWF: TOP
CFP + UWF: DeepDRiD
L3
Tasks: 4
Subtasks: 39
15,6067,237What are the positions of the haemorrhages in the fundus image?
A. Nasal to the optic disc center
B. Temporal to the optic disc center
C. Superior to the optic disc center
D. Inferior to the optic disc center
E. Not observed in the image
CFP: DDR, IDRiD, Retinal-Lesions
OCT: RETOUCH
L4
Tasks: 4
20,1775,114Which abnormalities can be seen in this fundus image?
A. Glaucoma
B. Diabetic retinopathy
C. No abnormality
D. Age-related macular degeneration
E. Hypertensive retinopathy
CFP: DDR, IDRiD, OIA-ODIR, JSIEC, RFMiD, Retinal-Lesions
OCT: NEH, OCTDL, OCTID, UCSD
UWF: TOP
CFP + OCT: MMC-AMD
CFP + UWF: DeepDRiD

4.3 Targeted Evaluation Modes

To assess an MLLM and its two key modules—the LLM and VE—the authors present three targeted evaluation modes (E-modes) in a bottom-up manner. 1 (Figure 1)

E-mode I: Linear-probe-based VE Evaluation

As illustrated in Figure 1b, linear probing trains a linear classification head for each task or subtask using the task-specific development dataset. 1

The authors omit tasks that cannot be directly addressed as classification problems, such as L3b, L3c, and L3d, as well as tasks that are trivial for linear probing, such as L1a and L1b.

E-mode II: Knowledge-prompted LLM Evaluation

Given a test image and its associated task-specific label, the authors convert the label into an indirect description by querying an expert knowledge database, EyeWiki. The description is then provided together with the image and question. 1 8

Example: L2b laterality recognition

Please choose the most suitable option based on the image, description, and question. Answer directly with the option letter.
Description: The fundus image shows the fovea located to the left of the optic disc.
Question: Which eye is shown in this image, the left or the right?
Options: A. Left eye; B. Right eye.

Illustrative example: L3a hard-exudate recognition

Please choose the most suitable option based on the image, description, and question. Answer directly with the option letter.
Description: The fundus image shows white or yellowish deposits with sharply defined margins.
Question: Does this fundus image show any evidence of hard exudates?
Options: A. No; B. Yes.

This example is included only to illustrate the prompt structure; it is not presented as a verbatim record from the released annotations.

E-mode III: Holistic Evaluation

This mode provides an end-to-end evaluation of the MLLM. 1

4.4 Performance Metrics

The paper calls its main metric “F1,” but the definition and implementation use the harmonic mean of sensitivity and specificity:

HSe,Sp=2 Se SpSe+Sp.H_{Se,Sp}=\frac{2\,Se\,Sp}{Se+Sp}.

This differs from the standard F1 score, which is the harmonic mean of precision and recall. For multiclass tasks, the implementation averages class-level HSe,SpH_{Se,Sp} scores and then aggregates hierarchically across subtasks, tasks, and levels. The result should therefore be described as the paper’s composite score, or HSe,SpH_{Se,Sp}, rather than conventional F1 or accuracy. 1 7

5. Evaluating MLLMs on FunBench

5.1 Choice of MLLMs

The authors select MLLMs at the 7B/8B scale, compiling a list of six general-purpose models and three medical models, as shown in Table 2. In addition, they include GPT-4o as a proprietary baseline and DINOv2-large as a VE baseline. 1

Table 2. Open-source MLLMs evaluated in the paper. Medical models are marked with *. 1 (Table 2)

MLLMHugging Face releaseVELLM
LLaVA-v1.5-7B2023.10CLIP-ViTVicuna-7B
*Qilin-Med-VL-Chat2023.12CLIP-ViTChinese-LLaMA2
LLaVA-v1.6-7B2024.01CLIP-ViTVicuna-7B
*LLaVA-Med-v1.5-7B2024.05CLIP-ViTMistral-7B
Qwen2-VL-7B2024.09Qwen2-ViTQwen2-7B
InternVL2.5-8B2024.12InternViTInternLM2.5-7B
*HuatuoGPT-Vision-7B2024.06CLIP-ViTQwen2-7B
Janus-Pro-7B2025.01ViT-SigLIPDeepSeek-LLM-7B
Qwen2.5-VL-7B2025.01Qwen2.5-ViTQwen2.5-7B

5.2 Results

VE Comparison

The performance of the different VEs is shown in the E-mode I section of Table 3. 1

The key results are as follows:

  1. Under the tested frozen-encoder linear-probe protocol and aggregate scoring procedure, DINOv2 is the best-performing VE, whereas CLIP-ViT has the lowest score among the evaluated encoders.
  2. The performance of the VEs is even worse than chance on L4b.

The authors state that these results suggest limitations in pure-vision solutions for fundus image analysis. I would restrict this conclusion to the tested visual representations and linear-probe protocol; it does not establish a limitation of nonlinear readouts, fine-tuned visual models, or all pure-vision approaches.

LLM Comparison

The LLM results are shown in the E-mode II section of Table 3. 1

The key results are as follows:

  1. Under the knowledge-prompted condition, systems such as InternVL2.5, HuatuoGPT-Vision, and Qwen can use label-derived ophthalmic descriptions to answer some fundus reading questions. Because this mode still includes the image and uses a description derived from the reference label, it does not provide a pure measurement of LLM knowledge.
  2. Their performance varies across tasks. For instance, the systems perform clearly better on L4c than on L4d. Both are disease-specific classification tasks, but their labels, data, and difficulty are not matched.
  3. Performance on L2b is close to chance.

The authors explain the second result by suggesting that DR-related materials are more abundant online than AMD-related materials, making LLMs more familiar with DR. For the poor performance on L2b, despite it being a simple task, the authors suggest that the skill may be too basic to be widely discussed, making relevant training data rare.

The authors conclude that the results suggest a fundamental limitation of the current data-driven paradigm: it produces powerful models that lack basic fundus reading skills.

MLLM Comparison

The performance of the MLLMs is summarized in the final section of Table 3. 1

  1. Among the tested open-source MLLMs, HuatuoGPT-Vision performs best, followed by Qwen2-VL and InternVL2.5. GPT-4o has the highest overall score when the proprietary baseline is included.
  2. HuatuoGPT-Vision and Qwen2-VL use LLMs with the same architecture, Qwen2-7B, yet the former’s VE (CLIP-ViT) is shown to be less effective than that of the latter (Qwen2-ViT). This is consistent with a possible benefit of domain-specific fine-tuning, but the comparison does not isolate that effect from differences in the visual encoder, connector, training data, or training procedure.
  3. As shown in Table 4, holistic model rankings are more strongly associated with rankings under the knowledge-prompted condition than with rankings from the VE linear-probe condition. This association does not establish the causal contribution of either module.

The authors state that the evaluated MLLMs have quite limited fundus reading skills. They also argue that the correlation analysis highlights the urgent need to develop a strong ophthalmic LLM.

Table 4. Correlation analysis based on mean-performance ranks. 1 (Table 4)

ModuleSpearman correlation with MLLM
LLM0.917
VE0.055

5.3 Conclusions

  • The evaluated MLLMs remain weak at fundus reading tasks involving anatomy perception, lesion analysis, and disease diagnosis.
  • The authors interpret the results as suggesting greater reliance on the LLM than on the VE; the reported correlations do not by themselves establish causal reliance.
  • HuatuoGPT-Vision’s performance is consistent with a possible benefit of domain-specific training, but the comparison does not isolate that factor.
  • Future training procedures need to consider the full range of fundus reading skills; otherwise, we risk developing an MLLM that lacks basic fundus reading abilities. 1

6. My Comments

The authors provide a reasonable benchmark for evaluating MLLMs’ fundus reading skills. We can adapt it to evaluate our models or agent workflows, and we can also extend its task-organization approach to other domains. The public code and annotations provide a useful starting point, although preparing the source images, adapting model interfaces, and defining tool access and inference budgets for agents still require additional work. 2 3

There are also several issues worth noting.

I could not identify an E-mode I feature-extraction and linear-probe training pipeline in the inspected version of the public repository (see issue #2). This observation concerns the inspected repository version and does not establish that no implementation exists elsewhere.

The authors’ explanations and hypotheses regarding the results are not sufficiently convincing.

For example, the authors suggest that the worse performance on AMD than on DR is related to the relative availability of training material. The observed performance difference motivates this hypothesis, but the paper provides no direct evidence about the relative coverage of DR and AMD in the models’ training data. Differences in image difficulty, class definitions, and question design remain possible alternatives.

The explanation that the simple laterality-recognition task (L2b) represents a rare skill is also not sufficiently convincing. One possible source of error is ambiguity in the coordinate or viewing convention: image-space left and right must be distinguished from anatomical laterality and from the patient’s perspective. Horizontal-flip augmentation may also affect orientation sensitivity, but this possibility requires an explicit controlled test; translation alone does not exchange left and right. Moreover, the relatively strong laterality results from the linear probe suggest that useful laterality information remains readable from the visual features, although the probe could be exploiting indirect cues. The error could therefore also arise in the connector, prompt interpretation, or answer mapping.

My main concern is the interpretation of E-mode II. The authors provide both an image and a description derived from the reference label, so the condition includes privileged ground-truth information. Such a condition can be useful diagnostically—for example, to ask whether the system can answer correctly when the relevant semantic evidence is supplied—but it is not a pure LLM evaluation. E-modes I and II also use different information, supervision, and task support, so comparing them cannot isolate the causal importance of the LLM and VE. The reported rank correlations are likewise associational rather than causal. 1 8

A description-only condition would help test whether the image continues to contribute once a label-derived description is supplied, although it would not provide a complete module decomposition. Additional controls could include blank, shuffled, or irrelevant images; descriptions written by experts without access to the reference label; and controlled module-replacement experiments. The authors’ concern that text-only input may not represent the behavior of an LLM after multimodal adaptation should also be acknowledged.

Although the paper provides useful targeted evaluations of VEs and LLMs, the separation is not complete. The end-to-end evaluation is more directly relevant to the intended image-based use case, but its stability would need to be established through repeated tests across random seeds, prompt variants, and input perturbations.

References

  1. Wei, Q., Qian, K., and Li, X. FunBench: Benchmarking Fundus Reading Skills of MLLMs. MICCAI 2025, LNCS 15965, pp. 278–288. DOI · Open-access manuscript.
  2. RUC AIMC Lab. FunBench source code and preparation instructions. Repository.
  3. AIMClab-RUC. FunBench VQA annotations. Hugging Face dataset.
  4. Hu, Y., et al. OmniMedVQA: A New Large-Scale Comprehensive Evaluation Benchmark for Medical LVLM. CVPR 2024. Official proceedings.
  5. Chen, P., et al. GMAI-MMBench: A Comprehensive Multimodal Evaluation Benchmark Towards General Medical AI. NeurIPS 2024. Official proceedings.
  6. Qin, Z., et al. LMOD: A Large Multimodal Ophthalmology Dataset and Benchmark for Large Vision-Language Models. Findings of NAACL 2025. ACL Anthology.
  7. RUC AIMC Lab. evaluate.py. Fixed-version source.
  8. RUC AIMC Lab. predict.py. Fixed-version source.
  9. Question about evaluation for emod1. FunBench GitHub issue #2. Issue.

Comments

Sign in — Sign in to join the conversation.

  1. …