Trustworthy GenAI Pipelines with Small Language Models
Keywords:
Enterprise GenAI Governance,Small Language Models (SLMs),Workflow Intelligence Automation,Governance-Aligned AI Pipelines,Responsible Enterprise AI,AI Compliance and Risk Management,Lightweight Language Models,Enterprise Workflow Orchestration,Secure AI Decision Systems,Explainable AI for Enterprises.Abstract
Governance-aligned Generative AI (GenAI) systems enable enterprise-grade GenAI capabilities by embedding governance safeguards directly into GenAI pipelines. Such pipelines instantiating enterprise workflows mitigate GenAI risks in GenAI model choices, foundation datasets, user prompts, and output consumption. Decisions regarding these GenAI pipeline elements should be made with the same formal structures and rigor applied to enterprise decision-making. A wealth of governance-aligned GenAI pipelines applied to enterprise processes would deliver comprehensive enterprise workflow intelligence at a fraction of the cost of traditional GenAI implementations.
For enterprises to harness Generative AI (GenAI) safely, enterprise governance objectives for GenAI must be realised. Candidates address the five enterprise corporate governance objectives—compliance, risk management, strategic alignment, performance and accountability—iteratively for all GenAI pathways; in computer science parlance, the task is to devise a breadth-first search through the enterprise GenAI pipeline graph. A foundational support system is a community-shared catalog of representative governance-aligned pipeline instances that address specific enterprise workflows or aspects thereof, deployed enterprise, domain or solution-wide with appropriate safeguards or controls. Candidate pipelines should instantiate enterprise workflows; the assets used in the pipeline or any GenAI input or output that may pose risk or compliance issues require particular attention. Enterprise workflows are risk scenarios against which GenAI adoption should be justified; pipeline adoption a risk-control measure.
References
1. Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., Brynjolfsson, E., et al. (2021). On the opportunities and risks of foundation models. arXiv.
2. Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., Lee, P., Lee, Y. T., Li, Y., Lundberg, S., Nori, H., & Zhang, Y. (2023). Sparks of artificial general intelligence: Early experiments with GPT-4. arXiv.
3. Radha, S., Gottimukkala, V. R. R., Thottara, S., Vandhana, K., & J, Gokulraj. (2025). Adaptive Video Streaming Over 5G Networks Using Deep Reinforcement Learning with Closed-Loop Feedback Mechanism for Bitrate Control. In 2025 International Conference on Communication, Computer, and Information Technology (IC3IT) (pp. 1–6). IEEE. 2025 International Conference on Communication, Computer, and Information Technology (IC3IT). https://doi.org/10.1109/ic3it66137.2025.11341184
4. Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al. (2022). PaLM: Scaling language modeling with pathways. Journal of Machine Learning Research.
5. Frantar, E., Ashkboos, S., Hoefler, T., & Alistarh, D. (2023). GPTQ: Accurate post-training quantization for generative pre-trained transformers. Proceedings of the International Conference on Learning Representations.
6. Mangalampalli, B. M., Bandi, V. D. V. K., Kolla, S. K., & Kumar, M. V. K. (2025). Towards Self-Evolving Healthcare Intelligence: Integrating Advanced Learning Systems with Real-Time Clinical Data Pipelines. Cultura: International Journal of Philosophy of Culture and Axiology, 22(12s), 464-486.
7. Guo, C., Zhao, Y., Yu, T., Zhang, C., & Xu, H. (2024). Trustworthy retrieval-augmented generation for large language models: A survey. arXiv.
8. Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, L., Wang, W., & Chen, Y. (2022). LoRA: Low-rank adaptation of large language models. Proceedings of the International Conference on Learning Representations.
9. Rani, P. S., Kummari, D. N., Yellanki, S. K., Meda, R., Koppolu, H. K. R., & Inala, R. (2025, July). Blockchain and AI for Securing Electrical Infrastructure. In 2025 2nd International Conference on Computing and Data Science (ICCDS) (pp. 1-6). IEEE.
10. Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., Ishii, E., Bang, Y., Madotto, A., & Fung, P. (2023). Survey of hallucination in natural language generation. ACM Computing Surveys, 55(12), 1–38.
11. Kandpal, N., Deng, H., Roberts, A., Wallace, E., & Raffel, C. (2023). Large language models struggle to learn long-tail knowledge. Proceedings of the International Conference on Machine Learning.
12. Lewis, P., Iyer, S., Minervini, P., Riedel, S., & Karpukhin, V. (2021). Retrieval-augmented generation for knowledge-intensive NLP tasks. Journal of Machine Learning Research, 22, 1–31.
13. Kummari, D. N., Burugulla, J. K. R., Malempati, M., Amistapuram, K., Garapati, R. S., & Nagabhyru, K. C. (2025, December). Enhancing Audit Compliance and Operational Efficiency in Manufacturing and Commercial Insurance Through Agentic AI and Data Engineering Frameworks. In 2025 IEEE International Conference on Communication Networks and Computing (CNC) (pp. 714-720). IEEE.
14. Li, X., Zhang, Y., Wang, S., & Liu, J. (2024). Efficient small language models for trustworthy generative AI: A survey. IEEE Access.
15. Lin, S., Hilton, J., & Evans, O. (2022). TruthfulQA: Measuring how models mimic human falsehoods. Proceedings of the Association for Computational Linguistics.
16. Nabende, P., & Wanyama, T. (2008). An expert system for diagnosing heavy-duty diesel engine faults. In Advances in computer and information sciences and engineering (pp. 384-389). Dordrecht: Springer Netherlands.
17. Liu, N. F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., & Liang, P. (2023). Lost in the middle: How language models use long contexts. Transactions of the Association for Computational Linguistics.
18. OpenAI. (2023). GPT-4 technical report. arXiv.
19. Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al. (2022). Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35.
20. Kolla, T. (2025). Generative AI for Intelligent Medical Coding and Healthcare Analytics. International Journal of Advanced Research in Computer Science & Technology (IJARCST), 8(6), 13285-13299.
21. Pan, S., Luo, L., Wang, Y., Chen, C., Wang, J., & Wu, X. (2023). Unifying large language models and knowledge graphs: A roadmap. IEEE Transactions on Knowledge and Data Engineering.
22. Penedo, G., Malartic, Q., Hesslow, D., Cojocaru, R., Cappelli, A., Alobeidli, H., & others. (2023). The RefinedWeb dataset for Falcon language models. Advances in Neural Information Processing Systems.
23. Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al. (2023). Llama 2: Open foundation and fine-tuned chat models. arXiv.
24. Mangalampalli, Bindu Madhavi, Sasi Kumar Kolla, Velangani Divya Vardhan Kumar Bandi, Uday Surendra Yandamuri, and PR Sudha Rani. "Designing intelligent healthcare ecosystems through adaptive data integration and autonomous learning systems." Vascular and Endovascular Review 8, no. 20s (2025): 330-347.
25. Wang, F., Zhang, Z., Zhang, X., Wu, Z., Wang, W., Li, R., Xu, J., Tang, X., He, Q., Ma, Y., Huang, M., & Wang, S. (2025). A comprehensive survey of small language models in the era of large language models: Techniques, enhancements, applications, collaboration with LLMs, and trustworthiness. ACM Transactions on Intelligent Systems and Technology.
26. Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q., & Zhou, D. (2022). Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems, 35.
27. Amistapuram, K., Pandiri, L., Raju, V. R., Paleti, S., Singireddy, S., & Sheelam, G. K. (2025). AI-Based Cloud Infrastructure and MLOps Frameworks for Scalable Data Engineering Across Banking and Insurance. In 2025 IEEE International Conference on Communication Networks and Computing (CNC) (pp. 186–192). IEEE. 2025 IEEE International Conference on Communication Networks and Computing (CNC). https://doi.org/10.1109/cnc68716.2025.11484532
28. Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., & Cao, Y. (2023). ReAct: Synergizing reasoning and acting in language models. Proceedings of the International Conference on Learning Representations.
29. Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al. (2023). GPT-4V(ision) system card. arXiv.
30. Kolla, S. H., & Mangala, N. (2025). DESIGNING AUTONOMOUS LLM AGENT FRAMEWORKS USING GEN AI PIPELINES TO ENHANCE CUSTOMER SERVICE MANAGEMENT AND KNOWLEDGE WORKFLOWS. Lex Localis-Journal of Local Self-Government, 23, 9719-9733.
31. Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., DasSarma, N., Drain, D., Ganguli, D., Henighan, T., Joseph, N., et al. (2022). Constitutional AI: Harmlessness from AI feedback. arXiv.
32. Chen, M., Tworek, J., Jun, H., Yuan, Q., Pinto, H. P., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., Ray, A., Puri, R., Krueger, G., Petrov, M., Khlaaf, H., Sastry, G., Mishkin, P., Chan, B., Gray, S., et al. (2021). Evaluating large language models trained on code. arXiv.
33. Sudhakar, A. V. V., Inala, R., Verma, A. K., Nag, K., Pandey, V., & Anand, P. S. (2025, September). Hybrid Rule-Based and Machine Learning Framework for Embedding Anti-Discrimination Law in Automated Decision Systems. In 2025 International Conference on Intelligent Communication Networks and Computational Techniques (ICICNCT) (pp. 1-6). IEEE.
34. Dettmers, T., Pagnoni, A., Holtzman, A., & Zettlemoyer, L. (2023). QLoRA: Efficient finetuning of quantized LLMs. Advances in Neural Information Processing Systems, 36.
35. Ding, N., Qin, Y., Yang, G., Wei, F., Yang, Z., Su, Y., Hu, S., Chen, Y., Chan, C. M., Chen, W., et al. (2023). Parameter-efficient fine-tuning of large-scale pre-trained language models: A survey. ACM Computing Surveys, 57(3), 1–40.
36. Kumar, I., Nagabhyru, K. C., IG, N., MV, P., & KV, S. (2025, October). Adaptive Meta-Knowledge Transfer Network with Feature Hallucination and Attention for Low-Shot Object Detection in Aerial Images. In 2025 International Conference on Communication, Computer, and Information Technology (IC3IT) (pp. 1-6). IEEE.
37. Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al. (2024). The Llama 3 herd of models. arXiv.
38. Gu, Y., Dong, L., Wei, F., & Huang, M. (2023). MiniLLM: Knowledge distillation of large language models. arXiv.
39. Singh, H., Bose, D., Nagubandi, A. R., Prabhu, S., & Naik, S. G. Cryptocurrency Market Spillovers: Risk Contagion Across Global Financial Systems.
40. Jiang, Z., Xu, F. F., Araki, J., & Neubig, G. (2023). How can we know when language models know? On the calibration of language models for question answering. Transactions of the Association for Computational Linguistics, 11, 962–977.
41. Kandpal, N., Wallace, E., & Raffel, C. (2023). Deduplicating training data makes language models better. Proceedings of the International Conference on Machine Learning.
42. Li, Y., Zhang, X., Wang, J., & Chen, H. (2024). Trustworthy generative AI: A survey of robustness, fairness, privacy, explainability, and accountability. IEEE Access, 12, 104221–104258.
43. Lin, B., Wang, B., Chen, X., & Liu, Z. (2024). Small language models: Survey, measurements, and applications. arXiv.
44. Loganathan, R. (2025). AGENTIC AI FRAMEWORKS FOR AUTONOMOUS RISK DETECTION AND COMPLIANCE REMEDIATION IN ENTERPRISE DATA CENTER OPERATIONS. Lex Localis-Journal of Local Self-Government, 23 (S6), 9672–9697.
45. Liu, H., Ning, R., Teng, Z., & Wang, Y. (2024). A survey on efficient large language models and small language models. ACM Computing Surveys.
46. Mialon, G., Dessì, R., Lomeli, M., Nalmpantis, C., Pasunuru, R., Raileanu, R., Rozière, B., Schick, T., Dwivedi-Yu, J., Celikyilmaz, A., Grave, E., LeCun, Y., & Scialom, T. (2023). Augmented language models: A survey. Transactions of the Association for Machine Learning Research.
47. Muennighoff, N., Tazi, N., Magne, L., & Reimers, N. (2023). MTEB: Massive text embedding benchmark. Proceedings of the European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases.
48. Nigam, N., Sireesha, B., Ediga, P., Segireddy, A. R., & Bokde, S. (2025, December). Comparative Evaluation of Cloud Security Algorithms Using Multiple Classifiers with an Optimized Intrusion Detection System. In 2025 IEEE 5th International Conference on ICT in Business Industry & Government (ICTBIG) (pp. 1-6). IEEE.
49. Nijkamp, E., Hayashi, H., Xie, T., Pang, B., Welleck, S., & Neubig, G. (2023). CodeGen2: Lessons for training LLMs on programming and natural languages. arXiv.
50. Qin, Y., Hu, S., Lin, Y., Chen, W., Yao, B., Zhou, X., Liang, S., Zhang, J., Xu, B., Zheng, J., et al. (2023). Tool learning with foundation models. ACM Computing Surveys, 56(9), 1–41.
51. Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., & Finn, C. (2024). Direct preference optimization: Your language model is secretly a reward model. Advances in Neural Information Processing Systems, 36.
52. Inala, R., Kaulwar, P. K., Nagabhyru, K. C., Adusupalli, B., & Arun Raj, S. R. (2025, October). Leveraging IEC 61850 for Interoperable and Resilient Smart Grid Communication Architecture. In International Conference on Microelectronics, Electromagnetics and Telecommunication (pp. 549-566). Cham: Springer Nature Switzerland.
53. Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M. A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al. (2023). LLaMA: Open and efficient foundation language models. Journal of Machine Learning Research, 24, 1–48.
54. Wang, L., Ma, C., Feng, W., Zhang, Y., Liu, H., Zhao, Y., Wang, H., Tang, H., Zhang, Y., et al. (2023). A survey on large language model acceleration based on efficient hardware and model compression. arXiv.
55. Amistapuram¹, K., Kolla, T., Bandi, V. D. V. K., Kolla⁴, S. K., & Rani, P. S. Journal of Rare Cardiovascular Diseases.
56. Zhao, W. X., Zhou, K., Li, J., Tang, T., Wang, X., Hou, Y., Min, Y., Zhang, B., Zhang, J., Dong, Z., et al. (2023). A survey of large language models. ACM Computing Surveys, 56(9), 1–44.
57. Anthropic. (2024). The Claude 3 model family: Opus, Sonnet, Haiku. arXiv.
58. Besta, M., Blach, N., Kubicek, A., Gerstenberger, R., Podstawski, M., Gianinazzi, L., Lehmann, T., Niewiadomski, H., Nyczyk, P., & Hoefler, T. (2024). Graph of Thoughts: Solving elaborate problems with large language models. Proceedings of the AAAI Conference on Artificial Intelligence.
59. Davuluri, P. N. AI-Augmented Sanctions Screening: Enhancing Accuracy and Latency in Real Time Compliance Systems.
60. Chen, X., Zhao, D., Wang, Y., Li, H., & Zhang, Y. (2024). Efficient inference techniques for small language models: A comprehensive survey. IEEE Access, 12, 122611–122639.
61. Dai, D., Sun, Y., Dong, L., Hao, Y., Sui, Z., & Wei, F. (2023). Why can GPT learn in-context? Language models implicitly perform gradient descent as meta-optimizers. Proceedings of the International Conference on Learning Representations.
62. Garapati, R. S., & Kanna, S. R. A Digital Twin‑Enabled Predictive Maintenance Framework Leveraging Multi‑Agent Reinforcement Learning and Industrial IoT Data.
63. Dong, Q., Li, L., Dai, D., Zheng, C., Ma, J., Li, R., Xia, H., Xu, J., Wu, Z., Chang, B., Sun, X., Xu, L., & Sui, Z. (2023). A survey on in-context learning. arXiv.
64. Gao, L., Biderman, S., Black, S., Golding, L., Hoppe, T., Foster, C., Phang, J., He, H., Thite, A., Nabeshima, N., Presser, S., & Leahy, C. (2021). The Pile: An 800GB dataset of diverse text for language modeling. arXiv.
65. Gunasekar, S., Zhang, Y., Aneja, J., Mendes, C., Del Giorno, A., Gopi, S., Javaheripi, M., Kauffmann, P., de Rosa, G., Saarikivi, O., et al. (2023). Textbooks are all you need. arXiv.
66. Izacard, G., & Grave, E. (2021). Leveraging passage retrieval with generative models for open-domain question answering. Proceedings of the European Chapter of the Association for Computational Linguistics, 874–880.
67. Ashokkumar, S., & Amistapuram, K. (2025, October). Attention-Guided Spatial Temporal Framework for Deepfake Detection on Social Video Platforms. In 2025 International Conference on Communication, Computer, and Information Technology (IC3IT) (pp. 1-6). IEEE.
68. Jiang, H., Ren, X., & Lin, B. Y. (2024). LLMs meet knowledge graphs: A survey of retrieval-augmented generation, reasoning, and trustworthiness. ACM Computing Surveys.
69. Kandpal, N., Deng, H., Roberts, A., Wallace, E., & Raffel, C. (2023). Large language models struggle to learn long-tail knowledge. Proceedings of the International Conference on Machine Learning.
70. Li, J., Xu, Y., Zhang, H., & Wang, X. (2024). Trustworthy artificial intelligence for foundation models: Challenges and opportunities. IEEE Intelligent Systems, 39(2), 18–29.
71. Reddy, V. A. R., & Kolla, S. K. (1984). Infrastructure-As-Code Practices For Regulated Healthcare Cloud Environments. Metallurgical and Materials Engineering, 30 (4), 1028–1042.
72. Liu, P., Yuan, W., Fu, J., Jiang, Z., Hayashi, H., & Neubig, G. (2023). Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing. ACM Computing Surveys, 55(9), 1–35.
73. Microsoft. (2024). Phi-3 technical report: A highly capable language model locally on your phone. arXiv.
74. Mangalampalli, B. M., & Kolla, S. K. (2025). Large Language Models for Automated Healthcare Data Dictionary Generation and Maintenance. Vascular and Endovascular Review, 8(20s), 363-375.
75. Mialon, G., Dessì, R., Nalmpantis, C., Lomeli, M., Pasunuru, R., Raileanu, R., Schick, T., Dwivedi-Yu, J., Celikyilmaz, A., Grave, E., LeCun, Y., & Scialom, T. (2023). Augmented language models: A survey. Transactions on Machine Learning Research.
76. Mistral AI. (2023). Mistral 7B. arXiv.
77. OpenAI. (2024). GPT-4o system card. arXiv.
78. Bandi, V. D. V. K. AI-Based Anomaly Detection Frameworks in Distributed Enterprise Data Systems.
79. Qin, Y., Hu, S., Lin, Y., Chen, W., Yao, B., Zhou, X., Liang, S., Zhang, J., Xu, B., Zheng, J., et al. (2023). Tool learning with foundation models. ACM Computing Surveys, 56(9), 1–41.
80. Sanh, V., Debut, L., Chaumond, J., & Wolf, T. (2021). Multitask prompted training enables zero-shot task generalization. Proceedings of the International Conference on Learning Representations.
81. Shumailov, I., Shumaylov, Z., Zhao, Y., Gal, Y., Papernot, N., & Anderson, R. (2024). AI models collapse when trained on recursively generated data. Nature, 631(8022), 755–759.
82. Team, G. (2024). Gemma: Open models based on Gemini research and technology. arXiv.
83. Abdin, M., Aneja, J., Bubeck, S., Eldan, R., Gunasekar, S., Harrison, M., He, J., Horvitz, E., Kauffmann, P., Lee, Y. T., Li, Y., et al. (2024). Phi-3 technical report: A highly capable language model for on-device AI. arXiv.
84. LEBCIR, I., Shah, C. A., & Appa Rao Nagubandi, D. S. M. D. FinTech and Financial Inclusion: Empirical Evidence from Emerging Markets.
85. Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. (2024). Advances in efficient language model deployment and inference. Journal of Machine Learning Research, 25, 1–36.
86. Cai, Z., Wang, Y., Zhang, H., Liu, X., & Chen, Y. (2024). Efficient small language models for edge intelligence: A survey. IEEE Internet of Things Journal, 11(18), 30214–30237.
87. Reddy, V. A. R. (2025). Journal of Rare Cardiovascular Diseases. Health, 5(3), 402-422.
88. Dao, T., Fu, D. Y., Ermon, S., Rudra, A., & Ré, C. (2022). FlashAttention: Fast and memory-efficient exact attention with IO-awareness. Advances in Neural Information Processing Systems, 35.
89. Dettmers, T., Lewis, M., Belkada, Y., & Zettlemoyer, L. (2024). Fine-tuning language models with quantization: Recent advances and challenges. ACM Computing Surveys, 57(5), 1–34.
90. Google DeepMind. (2024). Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. arXiv.
91. Huang, J., Chang, K. C. C., Wang, X., & Li, Y. (2024). Benchmarking trustworthy large and small language models: Robustness, factuality, and safety. IEEE Transactions on Artificial Intelligence, 5(4), 1881–1896.
92. GARAPATI, R. S. SYNERGETIC INTELLIGENCE Converging AI, Cloud, IoT, and Smart Automation for Real-Time Futures. CANEDA GLOBAL JOURNAL GROUP.
93. Liu, Z., Wang, Y., Chen, H., & Xu, J. (2025). Trustworthy small language models for secure and efficient generative AI systems: A comprehensive survey. Information Fusion, 106, 102514.
94. Wang, F., Zhang, Z., Zhang, X., Wu, Z., Wang, W., Li, R., Xu, J., Tang, X., He, Q., Ma, Y., Huang, M., & Wang, S. (2025). A comprehensive survey of small language models in the era of large language models: Techniques, enhancements, applications, collaboration with LLMs, and trustworthiness. ACM Transactions on Intelligent Systems and Technology.