CITAE: A Passage-Verifiable Retrieval-Augmented Generation Platform for Researchers and Thesis Writers

Abstract

Assistants based on large language models accelerate scholarly reading, but they hallucinate, which is unacceptable in academic work where every claim must be traceable to a source. We present CITAE, an open-source web application that integrates the complete scientific-literature workflow—discovery, read- ing, verification, organisation and citation—into a single system, and makes AI grounding visible to the user rather than hidden inside the model. Built as a React frontend over a layered Node.js/Express backend and a PostgreSQL database, its reading assistant ranks the user’s papers and highlights with Okapi BM25 (k1 = 1.5, b = 0.75) and injects the top-ranked, numbered passages into the prompt; the answer’s inline [n] markers are then bound back to the passages they cite and rendered as trust cards that expose the lit- eral supporting text and its source. A second, single-document assistant additionally verifies every quoted passage against the source and discards any that cannot be matched, so fabricated evidence never reaches the reader. A provider-agnostic client routes every model call through each user’s own keys, failing over across Groq, Google Gemini and OpenRouter, so the AI features run at no cost under a bring-your-own-key scheme and degrade gracefully. Around this core, CITAE unifies multi-source discovery over Crossref, Se- mantic Scholar, OpenAlex and arXiv with claim verification, paper comparison, literature reports, spaced review and citation in seven styles. We report an engineering evaluation on the production code: in-memory BM25 retrieval takes 0.5 ms over 100 and 6.4 ms over 1,000 passages; the deterministic verification guard attains precision, recall and F1 of 1.000 on a labelled 50-quotation probe set, rejecting all adversarial “genuine-prefix + fabricated-tail” strings; and an 86-case test suite covers the deterministic core. Offered bilingually with a Spanish-first default, free and open source, CITAE is an equity-oriented alternative to today’s fragmented, mostly proprietary and English-first landscape.

References

Audrin, C. and Audrin, B. (2022) Key factors in digital literacy in learning and education: a systematic literature review using text mining, Education and Information Technologies, 27(6), pp. 7395–7419. https://doi.org/10.1007/s10639-021-10832-5

Bangor, A., Kortum, P. T. and Miller, J. T. (2008) An empirical evaluation of the System Usability Scale, International Journal of Human–Computer Interaction, 24(6), pp. 574–594. https://doi.org/10.1080/10447310802205776

Bender, E. M., Gebru, T., McMillan-Major, A. and Shmitchell, S. (2021) On the dangers of stochastic parrots: can language models be too big?, Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, ACM, pp. 610–623. https://doi.org/10.1145/3442188.3445922

Brooke, J. (1996) SUS: a `quick and dirty' usability scale, in Usability Evaluation in Industry, Taylor & Francis, pp. 189–194. https://doi.org/10.1201/9781498710411-35

Cotton, D. R. E., Cotton, P. A. and Shipway, J. R. (2023) Chatting and cheating: ensuring academic integrity in the era of ChatGPT, Innovations in Education and Teaching International, 61(2), pp. 228–239. https://doi.org/10.1080/14703297.2023.2190148

de la Torre, A. and Baldeon-Calisto, M. (2024) Generative artificial intelligence in Latin American higher education: a systematic literature review, 2024 12th International Symposium on Digital Forensics and Security (ISDFS), IEEE, pp. 1–7. https://doi.org/10.1109/ISDFS60797.2024.10527283

Farias-Gaytan, S., Aguaded, I. and Ramirez-Montoya, M.-S. (2023) Digital transformation and digital literacy in the context of complexity within higher education institutions: a systematic literature review, Humanities and Social Sciences Communications, 10(1), article 386. https://doi.org/10.1057/s41599-023-01875-9

Featherstone, R., Walter, M., MacDougall, D., Morenz, E., Bailey, S., Butcher, R., Ford, C., Loshak, H. and Kaunelis, D. (2025) Artificial intelligence search tools for evidence synthesis: comparative analysis and implementation recommendations, Cochrane Evidence Synthesis and Methods, 3(5), e70045. https://doi.org/10.1002/cesm.70045

Flores-Vivar, J.-M. and Garcia-Penalvo, F.-J. (2023) Reflections on the ethics, potential and challenges of artificial intelligence in the framework of quality education (SDG4), Comunicar, 31(74), pp. 37–47. https://doi.org/10.3916/C74-2023-03

Gao, L., Dai, Z., Pasupat, P., Chen, A., Chaganty, A. T., Fan, Y., Zhao, V., Lao, N., Lee, H., Juan, D.-C. and Guu, K. (2023) RARR: researching and revising what language models say, using language models, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL, pp. 16477–16508. https://doi.org/10.18653/v1/2023.acl-long.910

Gao, Y., Xiong, Y., Gao, X., Jia, K., Pan, J., Bi, Y., Dai, Y., Sun, J., Wang, M. and Wang, H. (2023) Retrieval-augmented generation for large language models: a survey, arXiv preprint arXiv:2312.10997. https://doi.org/10.48550/arXiv.2312.10997

Garcia Aretio, L. (2019) The need for a digital education in a digital world, RIED. Revista Iberoamericana de Educacion a Distancia, 22(2), pp. 9–22. https://doi.org/10.5944/ried.22.2.23911

Garcia Aretio, L. (2021) COVID-19 and digital distance education: pre-confinement, confinement and post-confinement, RIED. Revista Iberoamericana de Educacion a Distancia, 24(1), pp. 9–32. https://doi.org/10.5944/ried.24.1.28080

Garcia-Martin, J. and Garcia-Sanchez, J.-N. (2022) The digital divide of know-how and use of digital technologies in higher education: the case of a college in Latin America in the COVID-19 era, International Journal of Environmental Research and Public Health, 19(6), article 3358. https://doi.org/10.3390/ijerph19063358

Hendricks, G., Tkaczyk, D., Lin, J. and Feeney, P. (2020) Crossref: the sustainable source of community-owned scholarly metadata, Quantitative Science Studies, 1(1), pp. 414–427. https://doi.org/10.1162/qss_a_00022

Hertzum, M. (2024) Concurrent or retrospective thinking aloud in usability tests: a meta-analytic review, ACM Transactions on Computer-Human Interaction, 31(3), pp. 1–29. https://doi.org/10.1145/3665327

Huang, L., Yu, W., Ma, W., Zhong, W., Feng, Z., Wang, H., Chen, Q., Peng, W., Feng, X., Qin, B. and Liu, T. (2025) A survey on hallucination in large language models: principles, taxonomy, challenges, and open questions, ACM Transactions on Information Systems, 43(2), pp. 1–55. https://doi.org/10.1145/3703155

Ince, D. C., Hatton, L. and Graham-Cumming, J. (2012) The case for open computer programs, Nature, 482(7386), pp. 485–488. https://doi.org/10.1038/nature10836

Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., Ishii, E., Bang, Y. J., Madotto, A. and Fung, P. (2023) Survey of hallucination in natural language generation, ACM Computing Surveys, 55(12), pp. 1–38. https://doi.org/10.1145/3571730

Jones, W. L. and Mastrorilli, T. (2022) Assessing the impact of an information literacy course on students' academic achievement: a mixed-methods study, Evidence Based Library and Information Practice, 17(2), pp. 61–87. https://doi.org/10.18438/eblip30090

Karpukhin, V., Oguz, B., Min, S., Lewis, P., Wu, L., Edunov, S., Chen, D. and Yih, W. (2020) Dense passage retrieval for open-domain question answering, Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), ACL, pp. 6769–6781. https://doi.org/10.18653/v1/2020.emnlp-main.550

Kasneci, E. et al. (2023) ChatGPT for good? On opportunities and challenges of large language models for education, Learning and Individual Differences, 103, article 102274. https://doi.org/10.1016/j.lindif.2023.102274

Kinney, R. et al. (2023) The Semantic Scholar open data platform, arXiv preprint arXiv:2301.10140. https://doi.org/10.48550/arXiv.2301.10140

Kratochvil, J. (2017) Comparison of the accuracy of bibliographical references generated for medical citation styles by EndNote, Mendeley, RefWorks and Zotero, The Journal of Academic Librarianship, 43(1), pp. 57–66. https://doi.org/10.1016/j.acalib.2016.09.001

Lewis, J. R. and Sauro, J. (2009) The factor structure of the System Usability Scale, in Human Centered Design, LNCS 5619, Springer, pp. 94–103. https://doi.org/10.1007/978-3-642-02806-9_12

Lewis, J. R. (2018) The System Usability Scale: past, present, and future, International Journal of Human–Computer Interaction, 34(7), pp. 577–590. https://doi.org/10.1080/10447318.2018.1455307

Lewis, P. et al. (2020) Retrieval-augmented generation for knowledge-intensive NLP tasks, Advances in Neural Information Processing Systems, 33, pp. 9459–9474. https://doi.org/10.48550/arXiv.2005.11401

Liu, N. F., Zhang, T. and Liang, P. (2023) Evaluating verifiability in generative search engines, Findings of the Association for Computational Linguistics: EMNLP 2023, ACL, pp. 7001–7025. https://doi.org/10.18653/v1/2023.findings-emnlp.467

Lund, B. D., Wang, T., Mannuru, N. R., Nie, B., Shimray, S. and Wang, Z. (2023) ChatGPT and a new academic reality: artificial intelligence-written research papers and the ethics of the large language models in scholarly publishing, Journal of the Association for Information Science and Technology, 74(5), pp. 570–581. https://doi.org/10.1002/asi.24750

Maynez, J., Narayan, S., Bohnet, B. and McDonald, R. (2020) On faithfulness and factuality in abstractive summarization, Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL, pp. 1906–1919. https://doi.org/10.18653/v1/2020.acl-main.173

Nature (2023) Tools such as ChatGPT threaten transparent science; here are our ground rules for their use (editorial), Nature, 613(7945), p. 612. https://doi.org/10.1038/d41586-023-00191-1

Nicholson, J. M., Mordaunt, M., Lopez, P., Uppala, A., Rosati, D., Rodrigues, N. P., Grabitz, P. and Rife, S. C. (2021) scite: a smart citation index that displays the context of citations and classifies their intent using deep learning, Quantitative Science Studies, 2(3), pp. 882–898. https://doi.org/10.1162/qss_a_00146

Nitsos, I., Malliari, A. and Chamouroudi, R. (2022) Use of reference management software among postgraduate students in Greece, Journal of Librarianship and Information Science, 54(1), pp. 95–107. https://doi.org/10.1177/0961000621996413

Nosek, B. A. et al. (2022) Replicability, robustness, and reproducibility in psychological science, Annual Review of Psychology, 73, pp. 719–748. https://doi.org/10.1146/annurev-psych-020821-114157

Priem, J., Piwowar, H. and Orr, R. (2022) OpenAlex: a fully-open index of scholarly works, authors, venues, institutions, and concepts, arXiv preprint arXiv:2205.01833. https://doi.org/10.48550/arXiv.2205.01833

Rashkin, H., Nikolaev, V., Lamm, M., Aroyo, L., Collins, M., Das, D., Petrov, S., Tomar, G. S., Turc, I. and Reitter, D. (2023) Measuring attribution in natural language generation models, Computational Linguistics, 49(4), pp. 777–840. https://doi.org/10.1162/coli_a_00486

Robertson, S. and Zaragoza, H. (2009) The probabilistic relevance framework: BM25 and beyond, Foundations and Trends in Information Retrieval, 3(4), pp. 333–389. https://doi.org/10.1561/1500000019

Scherbakov, D., Hubig, N., Jansari, V., Bakumenko, A. and Lenert, L. A. (2025) The emergence of large language models as tools in literature reviews: a large language model-assisted systematic review, Journal of the American Medical Informatics Association, 32(6), pp. 1071–1086. https://doi.org/10.1093/jamia/ocaf063

Valdivieso, T. and González, O. (2025) Generative AI tools in Salvadoran higher education: balancing equity, ethics, and knowledge management in the Global South, Education Sciences, 15(2), 214. https://doi.org/10.3390/educsci15020214

Wang, C. (2024) Exploring students' generative AI-assisted writing processes: perceptions and experiences from native and nonnative English speakers, Technology, Knowledge and Learning, 30(3), pp. 1825–1846. https://doi.org/10.1007/s10758-024-09744-3

Wilson, G., Bryan, J., Cranston, K., Kitzes, J., Nederbragt, L. and Teal, T. K. (2017) Good enough practices in scientific computing, PLOS Computational Biology, 13(6), e1005510. https://doi.org/10.1371/journal.pcbi.1005510

Yan, L., Sha, L., Zhao, L., Li, Y., Martinez-Maldonado, R., Chen, G., Li, X., Jin, Y. and Gasevic, D. (2024) Practical and ethical challenges of large language models in education: a systematic scoping review, British Journal of Educational Technology, 55(1), pp. 90–112. https://doi.org/10.1111/bjet.13370

Authors

DOI:

https://doi.org/10.31449/inf.v50i15.15471

Keywords:

academic literature tools, retrieval-augmented generation, BM25 passage retrieval, grounded language models, reference management, open-source scientific software, educational technology

Downloads

Published

09/09/2026

How to Cite

Vilca-Solorzano, R. A., Yana-Yucra, D. M., Torres-Cruz, F., Mamani-Calisaya, M. V., Herrera-Urtiaga, A. P., & Ibañez-Quispe, V. (2026). CITAE: A Passage-Verifiable Retrieval-Augmented Generation Platform for Researchers and Thesis Writers. Informatica, 50(15). https://doi.org/10.31449/inf.v50i15.15471