Solving Geometry Problems: A Text–Formula–Image Multimodal Parsing and Fusion Model

Solving geometry problems is a critical challenge in education, for it demands the integration of textual semantic descriptions, mathematical formula logic and spatial graphical information, as well as rigorous geometric theorem application and stepwise logical deduction. These are core capabilities that underpin the realization of personalized intelligent tutoring and efficient educational resource allocation. Traditional geometry problem solving methods often suffer from deficiencies in accuracy and the fusion of text, formula and image features. Hence, this paper proposes a method of solving geometry problems based on a text–formula–image (TFI) multimodal parsing and fusion model. The TFI parser employs a self-attention multilayer Transformer to enhance the extraction of logical relations among geometric text expressions. Meanwhile, it parses formulas into tree structures to overcome the loss of formula structural features, which utilizes symbolic embedding and tree-structured encoding to preserve hierarchical logical information and yields unified formula representations via a multi-granularity fusion module. The TFI parser also leverages a Feature Pyramid Network (FPN) for the accurate detection of geometric and non-geometric instances, resolves the issues of blurred segmentation for slender geometric elements and the inaccurate localization of small-sized symbols through mask averaging and RoIAlign, and generates high-dimensional image features using DenseNet-121. The TFI multimodal fusion model integrates a contrastive learning mechanism and constructs fused feature representations by stacking self-attention and cross-attention layers. This design effectively narrows the semantic gap between text, formula, and image features, addressing the inadequacy of traditional fusion approaches in deep cross-modal feature alignment. An attention-augmented Gated Recurrent Unit (GRU) network processes the fused TFI features to produce target operation trees and geometry solutions, ensuring interpretable and precise reasoning performance. The proposed method is evaluated on the PGDP5K and GeoEval datasets, and it achieves an average accuracy of 59.63% in geometry problem solving, which validates its effectiveness. This paradigm offers a viable technical approach for uniformly modeling complex educational tasks, including geometry problem solving and timetable scheduling.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC