Retrieval-Augmented Reliability-Aware Selective Inference for Visual Classification
Journal of Engineering Research and Sciences, Volume 5, Issue 9, Page # 1-14, 2026; DOI: 10.55708/js0509001
Keywords: Multimodal large language models, visual classification, selective prediction, retrievalaugmented, inference, reliability estimation, uncertainty estimation, abstention.
(This article belongs to the Special Issue on SP8 (Special Issue on Digital and Engineering Transformations in Science and Technology (SI-DETST-26)))
Export Citations
Cite
Hariharan, P. , Xu, H. and Yan, D. (2026). Retrieval-Augmented Reliability-Aware Selective Inference for Visual Classification. Journal of Engineering Research and Sciences, 5(9), 1ā14. https://doi.org/10.55708/js0509001
Pratheswaran Hariharan, Haiping Xu and Donghui Yan. "Retrieval-Augmented Reliability-Aware Selective Inference for Visual Classification." Journal of Engineering Research and Sciences 5, no. 9 (September 2026): 1ā14. https://doi.org/10.55708/js0509001
P. Hariharan, H. Xu and D. Yan, "Retrieval-Augmented Reliability-Aware Selective Inference for Visual Classification," Journal of Engineering Research and Sciences, vol. 5, no. 9, pp. 1ā14, Sep. 2026, doi: 10.55708/js0509001.
Multimodal large language models (MLLMs) can generate fluent visual responses even when the underlying visual prediction is weak, ambiguous, or incorrect. This work presents a retrieval-augmented, reliability-aware selective inference method that evaluates the strength and consistency of visual evidence before a prediction is communicated through a downstream multimodal response. A pretrained ResNet-50 encoder extracts normalized visual embeddings, and FAISS retrieves the topš reference images from an ImageNet-100 evidence database. Prediction reliability is assessed using retrieval similarity, class-support agreement, evidence margin, entropy-based uncertainty, and an aggregate reliability score. A decision gate then determines whether the prediction should be accepted, presented cautiously, or rejected through abstention or fallback. The selected decision is used to control the final user-facing response generated by the model. Experiments on ImageNet-100 show that the proposed method improves the accuracy of returned predictions from 85.84% to 88.88% at 89.04% coverage. The accepted error rate decreases from 14.16% to 11.12%, corresponding to a 3.04-percentage-point absolute reduction and a 21.48% relative reduction. Expected Calibration Error decreases from 7.40% to 5.87%, while high-reliability wrong predictions decrease from 263 to 226. These results demonstrate the potential of retrieval-derived evidence and selective decision gating as a post-hoc reliability mechanism for controlling visual predictions before they are incorporated into multimodal responses. The current evaluation is conducted in a controlled visual-classification setting and does not constitute a complete assessment of free-form MLLM hallucination.
- A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, I. Sutskever, āLearning transferable visual models from natural language supervisionā, āProceedings of the 38th International Conference on Machine Learningā, vol. 139 of Proceedings of Machine Learning Research, pp. 8748ā8763, PMLR, 2021.
- H. Liu, C. Li, Q. Wu, Y. J. Lee, āVisual instruction tuningā, āAdvances in Neural Information Processing Systemsā, vol. 36, pp. 34892ā34916, Curran Associates, Inc., 2023, doi: 10.52202/075280-1516.
- J. Li, D. Li, S. Savarese, S. Hoi, āBLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language modelsā, āProceedings of the 40th International Conference on Machine Learningā, vol. 202 of Proceedings of Machine Learning Research, pp. 19730ā19742, PMLR, 2023.
- W. Dai, J. Li, D. Li, A. M. H. Tiong, J. Zhao, W. Wang, B. Li, P. Fung, S. C. H. Hoi, āInstructBLIP: Towards general-purpose vision-language models with instruction tuningā, āAdvances in Neural Information Processing Systemsā, vol. 36, pp. 49250ā49267, Curran Associates, Inc., 2023, doi: 10.52202/075280-2142.
- J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds, R. Ring, E. Rutherford, S. Cabi, T. Han, Z. Gong, S. Samangooei, M. Monteiro, J. L. Menick, S. Borgeaud, A. Brock, A. Nematzadeh, S. Sharifzadeh, M. BiÅkowski, R. Barreira, O. Vinyals, A. Zisserman, K. Simonyan, āFlamingo: A visual language model for few-shot learningā, āAdvances in Neural Information Processing Systemsā, vol. 35, pp. 23716ā23736, Curran Associates, Inc., 2022, doi: 10.52202/068431-1723.
- D. Zhu, J. Chen, X. Shen, X. Li, M. Elhoseiny, āMiniGPT-4: Enhancing vision-language understanding with advanced large language modelsā, āInternational Conference on Learning Representationsā, 2024.
- A. Rohrbach, L. A. Hendricks, K. Burns, T. Darrell, K. Saenko, āObject hallucination in image captioningā, āProceedings of the 2018 Conference on Empirical Methods in Natural Language Processingā, pp. 4035ā4045, Association for Computational Linguistics, Brussels, Belgium, 2018, doi: 10.18653/v1/D18-1437.
- P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. Küttler, M. Lewis, W.-t. Yih, T. RocktƤschel, S. Riedel, D. Kiela, āRetrieval-augmented generation for knowledge-intensive NLP tasksā, āAdvances in Neural Information Processing Systemsā, vol. 33, pp. 9459ā9474, Curran Associates, Inc., 2020.
- K. Guu, K. Lee, Z. Tung, P. Pasupat, M. Chang, āRetrieval augmented language model pre-trainingā, āProceedings of the 37th International Conference on Machine Learningā, vol. 119 of Proceedings of Machine Learning Research, pp. 3929ā3938, PMLR, 2020.
- N. Papernot, P. McDaniel, āDeep k-nearest neighbors: Towards confident, interpretable and robust deep learningā, arXiv preprint arXiv:1803.04765, 2018, doi: 10.48550/arXiv.1803.04765.
- J. Johnson, M. Douze, H. JĆ©gou, āBillion-scale similarity search with GPUsā, IEEE Transactions on Big Data, vol. 7, no. 3, pp. 535ā547, 2021, doi: 10.1109/TBDATA.2019.2921572.
- D. Yan, Y. Wang, J. Wang, H. Wang, Z. Li, āK-nearest neighbor search by random projection forestsā, IEEE Transactions on Big Data, vol. 7, no. 1, pp. 147ā157, 2021, doi: 10.1109/TBDATA.2019.2908178.
- C. Guo, G. Pleiss, Y. Sun, K. Q. Weinberger, āOn calibration of modern neural networksā, āProceedings of the 34th International Conference on Machine Learningā, vol. 70 of Proceedings of Machine Learning Research, pp. 1321ā1330, PMLR, 2017.
- Y. Gal, Z. Ghahramani, āDropout as a Bayesian approximation: Representing model uncertainty in deep learningā, āProceedings of the 33rd International Conference on Machine Learningā, vol. 48 of Proceedings of Machine Learning Research, pp. 1050ā1059, PMLR, 2016.
- B. Lakshminarayanan, A. Pritzel, C. Blundell, āSimple and scalable predictive uncertainty estimation using deep ensemblesā, āAdvances in Neural Information Processing Systemsā, vol. 30, pp. 6402ā6413, Curran Associates, Inc., 2017.
- Y. Ovadia, E. Fertig, J. Ren, Z. Nado, D. Sculley, S. Nowozin, J. V. Dillon, B. Lakshminarayanan, J. Snoek, āCan you trust your modelās uncertainty? Evaluating predictive uncertainty under dataset shiftā, āAdvances in Neural Information Processing Systemsā, vol. 32, pp. 13991ā14002, Curran Associates, Inc., 2019.
- Y. Geifman, R. El-Yaniv, āSelective classification for deep neural networksā, āAdvances in Neural Information Processing Systemsā, vol. 30, pp. 4878ā4887, Curran Associates, Inc., 2017.
- D. Yan, P. Wang, M. Linden, B. S. Knudsen, T. W. Randolph, āStatistical methods for tissue array imagesāalgorithmic scoring and co-trainingā, The Annals of Applied Statistics, vol. 6, no. 3, pp. 1280ā1305, 2012, doi: 10.1214/12-AOAS543.
No related articles were found.
