GAST: Geometry-aware self-supervised transformer for efficient chest x-ray image retrieval
DOI:
https://doi.org/10.18488/76.v13i4.5201Keywords:
Attention mechanism, Chest X-ray, Deep learning, FAISS, Image retrieval, Medical imaging, Self-supervised learning, Vision transformer.Abstract
Chest radiography is currently one of the most widely used diagnostic techniques in the medical field. However, rapid and accurate retrieval of clinically correct cases from large-sized image databases is still a major challenge. The existing Content-Based Medical Image Retrieval (CBMIR) methods are mainly based on handcrafted features or hybrid optimization approaches. They face problems such as being weak in generalization, having high computational cost, and not being able to correctly identify important anatomical structures. To overcome these challenges, this study introduces a new framework called Geometry-Aware Self-Supervised Transformer (GAST). The recommended approach is based on a Vision Transformer backbone that is first trained using self-supervised methods such as masked autoencoding and contrastive learning. In addition, a geometry-aware attention module is introduced to guide the model toward important anatomical regions, particularly the lungs and heart. The proposed procedure was tested on a chest X-ray dataset available to the public. As per the results, GAST showed better performance compared to conventional deep learning models and achieved 98.85% accuracy. Also, there has been a significant improvement in retrieval criteria like Recall@K and Mean Average Precision. Faster response time was also achieved by using FAISS-based indexing. Overall, the research demonstrated that both interpretability and efficiency of the model were improved by the combination of geometry-aware attention and sparse feature learning. The proposed system will provide a reliable and beneficial solution in medical image retrieval and is likely to help in real-life clinical decision-making in the future.
