DataVista AI: An AI-Powered Data Catalog and Intelligence Copilot for Automated Metadata Management
Abstract
The exponential growth of heterogeneous enterprise data spread across CSV files, JSON documents, SQLite databases, and PostgreSQL instances has made it increasingly difficult for organizations to know what data they own, where it resides, and whether it can be trusted. DataVista AI is an AI-powered data catalog and intelligence copilot that automatically crawls structured data sources to construct a searchable metadata catalog. The system extracts schemas, computes statistical health profiles, discovers table-level relationships, generates business glossaries, and exposes an interactive AI copilot capable of conversational Text-to-SQL generation. A dual-engine language model manager routes requests to a cloud primary model with automatic failover to a secondary provider, while an Enterprise Privacy Mode and PII scrubbing layer ensure that raw sensitive values never leave the trust boundary. Built on a FastAPI backend, a ten-page Streamlit frontend, a lightweight Model Context Protocol (MCP) tool registry, and a LangChain-driven agent loop, DataVista AI demonstrates that automated, explainable data governance can be delivered without a heavyweight enterprise catalog platform.
References
Lewis P., Perez E., Piktus A., Petroni F., Karpukhin V., Goyal N., Küttler H., Lewis M., Yih W., Rocktäschel T., Riedel S., Kiela D. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. Advances in Neural Information Processing Systems. 2020, 33, 9459-9474p.
Reimers N., Gurevych I. Sentence-BERT: Sentence Embeddings Using Siamese BERT-Networks. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2019, 3982-3992p.
Yao S., Zhao J., Yu D., Du N., Shafran I., Narasimhan K., Cao Y. ReAct: Synergizing Reasoning and Acting in Language Models. International Conference on Learning Representations (ICLR). 2023.
Shi L. et al. A Survey on Employing Large Language Models for Text-to-SQL Tasks. ACM Computing Surveys. 2025.
Hong Z., Yuan Z., Zhang Q., Chen H., Dong J., Huang F., Huang X. Next-Generation Database Interfaces: A Survey of LLM-Based Text-to-SQL. IEEE Transactions on Knowledge and Data Engineering. 2025.
Busch J., Wu M. Automated Metadata Annotation: What Is and Is Not Possible with Machine Learning. Data Intelligence. 2023, 5(1), 122-138p.
Seetala S.R., Tran. Intelligent Data Catalogs Using Metadata Automation: Architectures, Standards, and Scalable Frameworks for Modern Data Ecosystems. International Journal of Scientific Research in Science and Technology. 2024, 11(3), 1037-1052p.
Singh M., Kumar A., Donaparthi S., Karambelkar G. Leveraging Retrieval Augmented Generative LLMs for Automated Metadata Description Generation to Enhance Data Catalogs. arXiv preprint arXiv:2503.09003. 2025.
Anthropic. Introducing the Model Context Protocol. Anthropic Research Publications. 2024.
Wu M. et al. Utilising a Large Language Model to Annotate Subject Metadata: A Case Study in an Australian National Research Data Catalogue. arXiv preprint arXiv:2310.11318. 2023.
Refbacks
- There are currently no refbacks.