Remove invalid image–instruction–reference records.
Visual Rubrics · Multimodal Post-Training
V-Rubrics: Visual Faithfulness via Rubric-Based Reinforcement Learning
*Equal Contributions ‡Project Lead
Abstract
Vision-language models can produce fluent answers that are visually wrong: a single unsupported object, chart value, or intermediate inference can invalidate an otherwise plausible response. We argue that this is a credit-assignment failure in multimodal post-training. Scalar outcome rewards say whether an answer is acceptable, but not which visual facts are grounded, which reasoning steps are valid, or which instruction constraints are missed. We introduce Visual Rubrics-Based Reinforcement Learning, which decomposes reference responses into atomic propositions and scores generated answers along Visual Faithfulness (VF), Reasoning Consistency (RC), and Instruction Following (IF). The resulting rubric items provide structured partial credit and localize rubric credit when supporting evidence spans are available. We first obtain a common SFT checkpoint by fine-tuning Qwen3-VL-8B-Instruct on the public OpenMMReasoner-SFT-874K corpus, adapting OpenMMReasoner’s cold-start data recipe. This checkpoint drives rejection sampling during the construction of V-Rubrics 50K, a 50K-example training set from 17 visually grounded sources, and initializes two GRPO variants using either scalar answer credit or component-wise, prefix-localized rubric credit. Rubric-based GRPO improves over both the shared SFT baseline and answer-level GRPO, with the largest gains on knowledge-oriented and visually grounded reasoning benchmarks. The results show rubrics as a useful reward abstraction for visual post-training.
Dataset
V-Rubrics 50K
50,248 rubric-annotated examples retained from 85,217 candidates across 17 public visual datasets.
Balance hard, medium, and simple curriculum bins from the rejection-sampling signal.
- 18,121
- hard · 36.1% · 0/8
- 25,306
- medium · 50.4% · 1/8–5/8
- 6,821
- simple · 13.6% · 6/8–8/8
Source-by-source counts: 85,217 selected candidates → 50,248 final
| Source dataset | Selected candidates* | Passed checks | Final | Final / selected |
|---|---|---|---|---|
| AI2D | 2,054 | 2,054 | 1,758 | 85.6% |
| ChartGalaxy | 2,900 | 1,500 | 1,289 | 44.4% |
| ChartQA | 760 | 760 | 687 | 90.4% |
| ChartQA-X | 4,145 | 1,500 | 1,314 | 31.7% |
| ChartX | 1,274 | 1,274 | 1,203 | 94.4% |
| CharXiv | 746 | 746 | 681 | 91.3% |
| DocVQA | 371 | 371 | 336 | 90.6% |
| Geometry3K | 1,074 | 1,074 | 919 | 85.6% |
| InfographicVQA | 4,034 | 2,500 | 2,071 | 51.3% |
| MathVerse | 2,942 | 2,200 | 1,290 | 43.8% |
| MM-K12 | 6,404 | 6,300 | 6,153 | 96.1% |
| PixMo-Count | 5,132 | 2,500 | 2,121 | 41.3% |
| ThinkLite-Hard | 9,226 | 9,226 | 8,624 | 93.5% |
| ThinkLite | 33,409 | 15,000 | 13,041 | 39.0% |
| ViRL39K | 7,304 | 7,304 | 6,188 | 84.7% |
| We-Math2.0-STD | 1,263 | 1,263 | 1,076 | 85.2% |
| We-Math2.0-Pro | 2,179 | 1,600 | 1,497 | 68.7% |
| Total | 85,217 | 57,172 | 50,248 | 59.0% |
* Candidate counts are pipeline selections, not full upstream dataset sizes.
Release contents and terms. The companion dataset packages the selected records, including image payloads, rubric annotations, and provenance metadata. Image and upstream-record rights remain source-dependent; inclusion in V-Rubrics does not relicense them, and the repository's Apache-2.0 license applies only to first-party software.
Our method
From visual data to rubric-guided RL
Performance
Performance Results
The full comparison reports all 22 models from the paper, with detailed panels for the five checkpoints evaluated under our setup.
general + knowledge
visual reasoning
67.74 → 68.04
Matched aggregate scores across all benchmark panels.
Overall average
Overall average
Shared panel axis: 55–70. Exact values are printed beside every bar.
MMBench-Dev and the knowledge-oriented MMMU family.
MMBench-Dev
MMMU Val
MMMU-Pro · 10c/Std
MMMU-Pro V
Knowledge Avg.Qwen3 Instruct 61.48Qwen3 Thinking 63.63SFT 58.31Answer GRPO 59.35Rubric GRPO 61.88
Shared panel axis: 50–90. Exact values are printed beside every bar.
Five visual-mathematical and grounded reasoning benchmarks.
MathVista mini
MathVision test
MathVerse V/O
DynaMath Worst
WeMath Loose
Math Avg.Qwen3 Instruct 59.50Qwen3 Thinking 63.50SFT 60.07Answer GRPO 63.19Rubric GRPO 63.63
Shared panel axis: 40–90. Exact values are printed beside every bar.
LogicVista and CharXiv reasoning.
LogicVista
CharXiv reasoning
Chart Avg.Qwen3 Instruct 56.20Qwen3 Thinking 59.02SFT 54.42Answer GRPO 58.81Rubric GRPO 59.51
Shared panel axis: 45–65. Exact values are printed beside every bar.
Select any paper metric to rank all 22 models; results not reported in the paper are shown as “-”.
Largest rubric gains: MMMU, MMMU-Pro, MathVision, DynaMath, WeMath Loose, and LogicVista. Answer GRPO remains higher: MMBench-Dev, MathVerse V/O, and CharXiv reasoning.
Examples
Eleven VQA records with all 84 generated rubrics shown in full.
Verified angle reasoning
Q Find the measure of ∠2 if m∠4 = m∠5.
Answer 53°
Manually verified. The image facts and 64° → 53° derivation are correct; m∠4 = m∠5 is extraneous to solving ∠2.
-
VF
VF_Triangle_IdentificationWeight +5Essential The response must identify from the image that the top-left triangle is composed of the interior angles ∠1, ∠2, and 63°.
-
VF
VF_Angles_On_Straight_LineWeight +5Essential The response must recognize from the image that the angle ∠1, the 69° angle, and the 47° angle are adjacent and lie on the same straight horizontal line.
-
VF
VF_Value_ExtractionWeight +4Important The response must accurately extract the precise numerical values from the diagram: 63° (top-left triangle), 69° (central intersection), and 47° (inside the top-right triangle).
-
VF
VF_Angle_69_ContextWeight +4Important The response must recognize that the 69° angle is situated in the space between the top-left and top-right triangles, rather than being an interior angle of either triangle.
-
VF
VF_Reading_ErrorsWeight -2Pitfall The response must not misread or hallucinate the numbers provided in the image (e.g., misreading 63° as 83° or 69° as 60°).
-
RC
RC_Angle_1_CalculationWeight +5Essential The response must logically deduce the measure of ∠1 by applying the supplementary angles rule for a straight line, calculating that ∠1 = 180° − 69° − 47° = 64°.
-
RC
RC_Angle_2_CalculationWeight +5Essential The response must logically calculate that ∠2 = 180° − 63° − 64° = 53° by applying the Triangle Angle Sum Theorem (180°) to the top-left triangle.
-
RC
RC_Extraneous_Information_HandlingWeight +4Important The response should solve for ∠2 without relying on or getting confused by the extraneous information “m∠4 = m∠5” provided in the prompt.
-
IF
IF_Final_AnswerWeight +5Essential The response must explicitly state that the final measure for ∠2 is 53 (or 53°).
Stacked-bar arithmetic
Q How much less CO₂ equivalent did PJSC Gazprom emit in 2020 than in 2019?
Answer 16.04 million metric tons
-
VF
Target_Company_IdentificationWeight +5Essential The response must correctly identify that 'PJSC Gazprom' corresponds to the blue segments at the base of the stacked bars, as indicated by the color legend at the bottom of the image.
-
VF
Year_SelectionWeight +5Essential The response must identify 'the previous year' relative to 2020 as being 2019, thus comparing data points from the 2019 and 2020 columns.
-
VF
Data_Extraction_2020Weight +4Important The response should identify the emission value for PJSC Gazprom (blue segment) in 2020 as 101.32 million metric tons of CO2 equivalent.
-
VF
Data_Extraction_2019Weight +4Important The response should identify the emission value for PJSC Gazprom (blue segment) in 2019 as 117.36 million metric tons of CO2 equivalent.
-
RC
Arithmetic_ReasoningWeight +4Important The response must correctly logically subtract the 2020 value from the 2019 value (117.36 - 101.32) to determine how much less the company emitted.
-
VF
Correct_Final_ValueWeight +5Essential The response must state the final answer is 16.04 (or approximately 16.04).
-
RC
Total_Value_ConfusionWeight -2Pitfall The response must not use the total emission numbers labeled at the top of the bars (236.45 for 2019 and 210.32 for 2020), which would result in an incorrect difference of 26.13.
-
IF
Unit_InclusionWeight +2Optional The response should ideally include the unit of measurement 'million metric tons of CO2 equivalent' as labeled on the y-axis.
Scientific plot reading
Q At σ = 2 in the ccf-auc plot, which method is closest to 0.6 accuracy?
Answer Mutual information
-
VF
Target_Plot_IdentificationWeight +5Essential The response must correctly identify and use the data from the middle plot, which is titled 'ccf-auc'.
-
VF
X_Axis_AlignmentWeight +5Essential The response must look at the data points corresponding to the value '2' on the x-axis (noise level (sigma)).
-
VF
Method_Data_Extraction_MIWeight +4Important The response should recognize that the 'Mutual information' method (indicated by the red line with left-pointing triangles) has an accuracy of approximately 0.65 at noise level 2.
-
VF
Method_Data_Extraction_OthersWeight +4Important The response should recognize that the other methods (Distance, Kendall, Pearson) cluster together at a higher accuracy of approximately 0.75-0.80 at noise level 2.
-
RC
Comparative_ReasoningWeight +4Important The response must logically conclude that since ~0.65 is closer to 0.6 than the values of the other methods (~0.75-0.80), 'Mutual information' is the correct answer.
-
IF
Final_Answer_SelectionWeight +5Essential The response must explicitly state that 'Mutual information' is the method with an accuracy closest to 0.6.
Document lookup
Q What postage covers all items on invoices 1176 and 1175?
Answer Aus $50
-
VF
Correct_Amount_ExtractionWeight +5Essential The response must state that the amount charged for postage is 50, as shown next to the text 'Plus postage to cover all items on invoice 1176 and 1175'.
-
VF
Currency_IdentificationWeight +3Important The response should include the currency unit 'Aus $' or 'Australian dollars' as printed in the receipt next to the value 50.
-
VF
Invoice_Number_ConfusionWeight -2Pitfall The response must not confuse the invoice numbers mentioned in the line (1176 or 1175) with the actual postage fee amount (50).
-
VF
Visual_Anchor_ValidationWeight +4Important The response should correctly reference the specific line on the receipt that reads 'Plus postage to cover all items on invoice 1176 and 1175' to justify the answer.
-
RC
Reasoning_Logical_MappingWeight +4Important The response must logically link the phrase 'postage to cover all items' from the question to the corresponding dollar value of 50 in the same row, rather than referencing the Sub Total (4,595) or the Total (4,645).
-
IF
Direct_Answer_FormatWeight +2Optional The response should provide a direct and concise answer to the specific question about the postage amount.
Inscribed-circle geometry
Q Circle P is inscribed in equilateral triangle LMN. What is its circumference?
Answer 8π/√3 in
-
VF
Side_Length_RecognitionWeight +5Essential The response correctly identifies that the side length of triangle LMN is 8 inches, as indicated by the label “8 in.” and the arrows pointing to segment LN in the image.
-
VF
Triangle_Type_IdentificationWeight +5Essential The response identifies triangle LMN as equilateral, consistent with both the prompt text and the visually regular blue triangle shown in the image.
-
VF
Inscribed_Circle_LogicWeight +4Important The response recognizes that circle P is inscribed within the triangle, meaning it is tangent to all three blue lines forming the triangle LMN in the image.
-
VF
Vertex_and_Center_LabelsWeight +2Optional The response correctly references the triangle vertices as L, M, and N and the circle center as P, following the black text labels in the diagram.
-
VF
Dimension_Label_MisinterpretationWeight -2Pitfall The response must not incorrectly assume the “8 in.” label refers to the diameter or radius of the circle P rather than the side length of the triangle.
-
RC
Inradius_CalculationWeight +4Important The response correctly calculates the inradius r of circle P from the side length s = 8 using the formula r = s / (2√3), resulting in r = 4/√3.
-
RC
Circumference_Calculation_ResultWeight +5Essential The response provides the correct final circumference value by applying C = 2πr, resulting in 8π/√3 (or 8√3π/3 or approximately 14.51 inches).
-
IF
Mathematical_Form_ConsistencyWeight +3Optional The response provides the answer in an exact symbolic form involving π and radicals, matching the style of the reference answer 8π/√3.
Function transformation
Q The logarithm g(x) transforms f(x) = ln x. State g(x).
Answer g(x) = ln x + 4
-
VF
Identify_f_FunctionWeight +5Essential The response must correctly identify that the base function shown in the image is f(x) = ln x.
-
VF
Coordinate_f_at_1Weight +3Important The response must correctly identify that the curve for f(x) passes through the point (1, 0) on the coordinate grid.
-
VF
Coordinate_g_at_1Weight +5Essential The response must correctly identify that the curve for g(x) passes through the point (1, 4) on the coordinate grid.
-
VF
Vertical_Asymptote_ObservationWeight +4Important The response must observe that both f(x) and g(x) have a vertical asymptote at x = 0, indicating no horizontal shift occurred.
-
RC
Vertical_Shift_ReasoningWeight +4Important The response must logically conclude that g(x) is a vertical translation of f(x) upward by 4 units based on the difference in y-values at x = 1 (4 − 0 = 4).
-
RC
Final_Equation_CorrectnessWeight +5Essential The response must state the final equation for g(x) as g(x) = ln x + 4.
-
IF
Function_NotationWeight +2Optional The response should explicitly use the notation g(x) = … when stating the final equation as requested by the prompt.
-
RC
Horizontal_Shift_ErrorWeight -2Pitfall The response must not incorrectly identify the transformation as a horizontal shift, such as writing g(x) = ln(x + 4) or g(x) = ln(x − 4).
Fine-grained counting
Q How many bowls or basins are in the image?
Answer 4
-
VF
Exact_Quantity_MatchWeight +5Essential The response must state that there are exactly 4 bowls or basins present in the image, matching the reference count.
-
VF
Identification_Green_Salad_BowlWeight +3Important The response should correctly identify the green, square-ish bowl in the upper right containing a salad with pomegranate seeds as one of the bowls.
-
VF
Identification_White_Potato_BowlWeight +3Important The response should correctly identify the white round bowl on the right side of the table containing mashed potatoes as one of the bowls.
-
VF
Identification_Red_Stuffing_BasinWeight +3Important The response should correctly identify the large red oval dish/basin on the left containing stuffing as one of the basins.
-
VF
Identification_Orange_Gravy_BoatWeight +3Important The response should correctly identify the orange gravy boat (a small basin) located in the center of the table.
-
RC
Logical_Exclusion_PlatterWeight +4Important The response must logically exclude the flat red platter at the bottom center of the image, which holds the turkey, as it does not qualify as a bowl or basin.
-
RC
Reasoning_Pie_Dish_ExclusionWeight +3Important The response should distinguish the bowl/basin-like containers from the shallow pie dish in the top left or the wire cooling rack it sits on to maintain the count of 4.
-
IF
Direct_Answer_FormatWeight +2Optional The response should ideally provide the number '4' directly or as the primary focus of the sentence to meet the user's information need efficiently.
-
VF
Hallucination_PitfallWeight -2Pitfall The response must not claim there are other bowls not visible in the image or count non-container items like spoons or placemats.
Halving sequence
Q A beaver starts with 64 butterflies; half leave after each photo, and the last photo has 2. How many photos?
Answer 6 photos · option A
-
VF
Initial_Count_ExtractionWeight +5Essential The response must correctly extract the value '64' from the image text as the number of butterflies present in the first photo.
-
VF
Final_Count_ExtractionWeight +5Essential The response must correctly extract the value '2' from the image text as the number of butterflies present in the last photo.
-
VF
Rule_FaithfulnessWeight +5Essential The response must accurately reflect the rule stated in the image text that 'half the butterflies fly away' after each photo is taken.
-
VF
Visual_Subject_ContextWeight +2Optional The response identifies the specific visual elements of the image, such as a beaver holding a blue and green camera photographing a group of pink, blue, and green butterflies.
-
RC
Mathematical_Halving_LogicWeight +4Important The response must demonstrate a logically consistent chain of halving the population (64, 32, 16, 8, 4, 2) based on the rule identified from the image.
-
RC
Deduction_of_Total_PhotosWeight +5Important The response must correctly conclude that the sequence of counts (64, 32, 16, 8, 4, 2) corresponds to a total of 6 photos taken.
-
IF
Selection_of_Correct_OptionWeight +5Essential The response must explicitly identify Choice A as the correct answer.
-
RC
Off_By_One_ErrorWeight -1Pitfall The response should not miscalculate the total number of photos as 5 or 32 by either excluding the first/last photo in the sequence or confusing the count with the number of butterflies remaining after the first photo.
Stacked-bar composition
Q How much more tax revenue than public spending did the USA report in 2021?
Answer $50 billion
-
VF
VF_USA_Public_SpendingWeight +5Essential The response must correctly identify from the chart that the 'Public Spending' for the USA (represented by the blue bar segment) is 400 billion.
-
VF
VF_USA_Total_HeightWeight +4Important The response must correctly identify that the total height of the stacked bar for the USA reaches 850 billion on the y-axis.
-
VF
VF_USA_Tax_Revenue_CalculationWeight +5Essential The response must correctly identify or calculate that the 'Tax Revenue' for the USA (represented by the orange bar segment) is 450 billion (total height 850 - public spending 400).
-
RC
RC_Surplus_CalculationWeight +5Essential The response must logically subtract the public spending (400 billion) from the tax revenue (450 billion) to determine the surplus.
-
VF
VF_Correct_Final_ValueWeight +5Essential The response must state the final answer is 50 billion.
-
VF
VF_Legend_InterpretationWeight +4Important The response must correctly interpret the legend, identifying blue as 'Public Spending' and orange as 'Tax Revenue'.
-
VF
Pitfall_Absolute_Height_MisreadingWeight -2Pitfall The response must not incorrectly claim that the tax revenue is 850 billion by misinterpreting the total stacked bar height as the value for the top segment only.
-
IF
IF_Unit_UsageWeight +2Optional The response should include the correct unit 'billion' or '$ billion' as specified in the y-axis label.
Food-art recognition
Q What animal is being mimicked in this food art?
Answer Dolphin
-
VF
Animal_IdentificationWeight +5Essential The response must explicitly identify the animal being mimicked as a dolphin.
-
VF
Visual_Medium_BananasWeight +4Important The response should acknowledge that the dolphins are crafted from yellow bananas, utilizing their curved shape and stems.
-
VF
Visual_Detail_Mouth_and_GrapeWeight +3Important The response should accurately note visual details such as the banana stems being split to look like open mouths, specifically highlighting that the tallest dolphin is holding a green grape in its mouth.
-
VF
Visual_Detail_EyesWeight +2Optional The response may mention that small dark dots (likely peppercorns, cloves, or ink) have been added to the bananas to represent the dolphins' eyes.
-
RC
Contextual_InterpretationWeight +3Important The response should logically interpret the scene as dolphins leaping or emerging from a body of water, represented by the bed of green grapes on the plate.
-
IF
Direct_AnswerWeight +5Essential The response must directly answer the question 'What animal is being mimicked' without irrelevant rambling or identifying the wrong creature (e.g., calling them birds or sharks).
Composite area
Q Find the total area of the plotted land region.
Answer 3 + π/2
-
VF
VF_Coordinates_IdentificationWeight +5Essential The response must correctly identify the coordinates of the key points from the image: A(-2,0), C(2,0), E(-1,1), D(1,1), and B(0,2).
-
VF
VF_Circle_PropertiesWeight +5Essential The response must identify that the circle is centered at point O(0,1) and has a radius of 1, as it passes through points I(0,0), D(1,1), B(0,2), and E(-1,1).
-
VF
VF_Plot_CompositionWeight +4Important The response should recognize that the 'entire plot' shown in the image is composed of a trapezoid ACDE (bottom half) and a semicircle EBD (top half).
-
RC
RC_Trapezoid_Area_CalculationWeight +4Important The response must correctly calculate the area of the lower trapezoid ACDE using the formula 0.5 * (base1 + base2) * height, where base1 (AC) = 4, base2 (ED) = 2, and height = 1, resulting in an area of 3.
-
RC
RC_Semicircle_Area_CalculationWeight +4Important The response must correctly calculate the area of the upper semicircle EBD using the formula 0.5 * π * r², where the radius r = 1, resulting in an area of π/2.
-
RC
RC_Final_SummationWeight +5Essential The response must sum the areas of the two parts (trapezoid and semicircle) to provide the final result of 3 + π/2.
-
RC
Pitfall_Triangle_OnlyWeight -2Pitfall The response must not state that the area is simply the area of triangle ABC (which is 4), as this ignores the circular boundary shown at the top of the plot.
-
IF
IF_Clear_ExpressionWeight +1Optional The response should clearly state the final numerical or symbolic expression (3 + π/2) as the answer to the question.
Labels show the source dataset and dataset-card license; individual media may carry additional terms.
Limitations
Structured feedback is still model-generated feedback
Generated rubrics may inherit answer bias, ambiguity, or unsupported assumptions.
Fuzzy span matching is not exact token-level supervision.
Benchmark accuracy only approximates visual faithfulness.
Qwen-family judges may favor Qwen-family policies.
Citation
BibTeX
If you use V-Rubrics, please cite:
@misc{tian2026vrubrics,
title = {V-Rubrics: Visual Faithfulness via Rubric-Based
Reinforcement Learning},
author = {Shulin Tian and Minglun Li and Yuhao Dong and
Hao Ding and Jiarui Yao and Haiwen Diao and
Jingkang Yang and Hongyuan Zhu and Ziwei Liu},
year = {2026},
url = {https://shulin16.github.io/v-rubrics/}
}