Bimodal network architectures for automatic generation of image annotation from text

Medical image analysis practitioners have embraced big data methodologies.\nThis has created a need for large annotated datasets. The source of big data is\ntypically large image collections and clinical reports recorded for these\nimages. In many cases, however, building algorithms aimed at segmentation and\ndetection of disease requires a training dataset with markings of the areas of\ninterest on the image that match with the described anomalies. This process of\nannotation is expensive and needs the involvement of clinicians. In this work\nwe propose two separate deep neural network architectures for automatic marking\nof a region of interest (ROI) on the image best representing a finding\nlocation, given a textual report or a set of keywords. One architecture\nconsists of LSTM and CNN components and is trained end to end with images,\nmatching text, and markings of ROIs for those images. The output layer\nestimates the coordinates of the vertices of a polygonal region. The second\narchitecture uses a network pre-trained on a large dataset of the same image\ntypes for learning feature representations of the findings of interest. We show\nthat for a variety of findings from chest X-ray images, both proposed\narchitectures learn to estimate the ROI, as validated by clinical annotations.\nThere is a clear advantage obtained from the architecture with pre-trained\nimaging network. The centroids of the ROIs marked by this network were on\naverage at a distance equivalent to 5.1% of the image width from the centroids\nof the ground truth ROIs.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC