Graph-Based Multimodal Contrastive Learning for Chart Question Answering

Chart question answering (ChartQA) is challenged by the heterogeneous composition of chart elements and the subtle data patterns they encode.This work introduces a novel joint multimodal scene graph framework that explicitly models the relationships among chart components and their underlying structures.The framework integrates both visual and textual graphs to capture structural and semantic characteristics, while a graph contrastive learning strategy aligns node representations across modalities-enabling their seamless incorporation into a transformer decoder as soft prompts.Moreover, a set of tailored Chain-of-Thought (CoT) prompts is proposed to enhance multimodal large language models (MLLMs) in zero-shot scenarios by mitigating hallucinations.Extensive evaluations on benchmarks including ChartQA, OpenCQA, and ChartX demonstrate significant performance improvements and validate the efficacy of the proposed approach.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC