Secure Edge AI for Financial Document Processing
Keywords:
Edge Cloud Ecosystem, Secure Feature Extraction, Financial Document Privacy, Generative AI Feature Synthesis, Agentic AI Supervision, Privacy Preserving Machine Learning, Edge Layer Intelligence, Cloud Layer Analytics, Decentralized AI Architecture, Sensitive Data Protection, Semantic Feature Enrichment, Downstream Task Accuracy, Synthetic Feature Generation, Private Label AI Models, Zero Data Leakage Design, Secure AI Pipelines, Confidential Document Processing, Trustworthy AI Systems, Edge AI Security, Federated Intelligence.Abstract
Enhancing security and privacy during feature extraction of financial documents is of paramount importance. Conventional concerns of data loss during feature extraction can be addressed in a natural way by placing feature extraction on the edge layer of an edge-cloud ecosystem. Recent developments in Generative and Agentic AI open opportunities for generating features for various tasks without sharing or leaking sensitive details about the data. Such capabilities can be expressed as Generative and Agentic AI-enhanced feature extraction—combining agentic capabilities of Generative AI to supervise the feature extraction process and also provide additional semantics and or representations that help improve the accuracy of the downstream task. This framework is shown to naturally extend to both edge-layer and cloud-layer processing. Security and privacy requirements during the feature-extraction stage are met by suitably applying standard security mechanisms without any changes to the underlying principle or operation of the feature extraction process.
Generative AI is an emerging technology with numerous applications in various domains. The recent capability of generating new data and labels from a few or a small set of samples but without sharing or leaking the private details in the original data has opened another important dimension. This capability not only allows the origin of the samples to remain secret, but also helps protect the sensitive nature of the data. Generative AI can thus accentuate the benefits of decentralization by further reducing the role and dominance of central-proxy-service providers in edge and cloud ecosystems. Instead of consuming and storing all the edge- and cloud-layer data, the major job can be reduced to merely providing the infrastructure for empowering other users in the ecosystem with the appropriate private-label Generation AI capability required for their workloads.
References
1. Appalaraju, S., Jasani, B., Kota, B. U., Xie, Y., & Manmatha, R. (2021). DocFormer: End-to-end transformer for document understanding. Proceedings of the IEEE/CVF International Conference on Computer Vision, 993–1003.
2. Arbreha, H. G., Hayajneh, M., & Serhani, M. A. (2022). Federated learning in edge computing: A systematic survey. Sensors, 22(2), 450.
3. Sriram, H. K., Challa, K., Gadi, A. L., & Singreddy, S. (2025). AI and Cloud-Driven Transformation in Finance, Insurance, and the Automotive Ecosystem: A Multi-Sectoral Framework for Credit Risk, Mobility Services, and Consumer Protection. Anil Lokesh and singreddy, Sneha, AI and Cloud-Driven Transformation in Finance, Insurance, and the Automotive Ecosystem: A Multi-Sectoral Framework for Credit Risk, Mobility Services, and Consumer Protection (March 15, 2025).
4. Brecko, A., Kajati, E., Koziorek, J., & Zolotova, I. (2022). Federated learning for edge computing: A survey. Applied Sciences, 12(18), 9124.
5. Chen, Z., Chen, W., Smiley, C., Shah, S., Borova, I., Langdon, D., Moussa, R., Beane, M., Huang, T.-H., Routledge, B., & Wang, W. Y. (2021). FinQA: A dataset of numerical reasoning over financial data. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, 3697–3711.
6. Hong, T., Kim, D., Ji, M., Hwang, W., Nam, D., & Park, S. (2022). BROS: A pre-trained language model focusing on text and layout for better key information extraction from documents. Proceedings of the AAAI Conference on Artificial Intelligence, 36(10), 10767–10775.
7. Seenu, A., Sheelam, G. K., Motamary, S., Meda, R., Koppolu, H. K. R., & Inala, R. (2025). AI-Driven Innovations in Infrastructure Management with 6G Technology. In 2025 2nd International Conference on Computing and Data Science (ICCDS) (pp. 1–6). IEEE. 2025 2nd International Conference on Computing and Data Science (ICCDS). https://doi.org/10.1109/iccds64403.2025.11209649
8. Huang, Y., Lv, T., Cui, L., Lu, Y., & Wei, F. (2022). LayoutLMv3: Pre-training for document AI with unified text and image masking. Proceedings of the 30th ACM International Conference on Multimedia, 4083–4091.
9. Lee, C.-Y., Li, C.-L., Dozat, T., Perot, V., Su, G., Hua, N., Ainslie, J., Wang, R., Fujii, Y., & Pfister, T. (2022). FormNet: Structural encoding beyond sequential modeling in form document information extraction. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics, 3735–3754.
10. Narasareddy Annapareddy, V., Pamisetty, A., Malempati, M., Kaulwar, P. K., & Bhardwaj Komaragiri, V. (2025, April). Enhancing Solar Power System Efficiency Through AI-Driven Predictive Maintenance and Cloud-Based Infrastructure Stability Solutions. In International Conference on Smart Computing and Informatics (pp. 328-337). Cham: Springer Nature Switzerland.
11. Lee, K., Joshi, M., Turc, I. R., Hu, H., Liu, F., Eisenschlos, J. M., Khandelwal, U., Shaw, P., Chang, M.-W., & Toutanova, K. (2023). Pix2Struct: Screenshot parsing as pretraining for visual language understanding. Proceedings of the 40th International Conference on Machine Learning, 202, 18893–18912.
12. Li, J., Xu, Y., Cui, L., & Wei, F. (2022). MarkupLM: Pre-training of text and markup language for visually rich document understanding. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics, 6078–6087.
13. Challa, S. R., Burugulla, J. K. R., Pamisetty, A., Challa, K., & Paleti, S. (2025, April). AI and ML-Powered Cybersecurity Strategies for Cloud Computing: Ensuring Infrastructure Stability in Financial and Retail Sectors. In International Conference on Smart Computing and Informatics (pp. 315-327). Cham: Springer Nature Switzerland.
14. Li, J., Xu, Y., Lv, T., Cui, L., Zhang, C., & Wei, F. (2022). DiT: Self-supervised pre-training for document image transformer. Proceedings of the 30th ACM International Conference on Multimedia, 3530–3539.
15. Li, Q., Wen, Z., Wu, Z., Hu, S., Wang, N., Li, Y., Liu, X., & He, B. (2023). A survey on federated learning systems: Vision, hype and reality for data privacy and protection. IEEE Transactions on Knowledge and Data Engineering, 35(4), 3347–3366.
16. Singireddy, S., Pandiri, L., Motamary, S., Kummari, D. N., Somu, B., & Lakkarasu, P. (2025, August). Blockchain-Powered Claims Validation for Enhancing Trust in Health and Auto Insurance Ecosystems. In International Conference on Artificial Intelligence: Theory and Applications (pp. 26-38). Cham: Springer Nature Switzerland.
17. Li, Y., Qian, Y., Yu, Y., Qin, X., Zhang, C., Liu, Y., Yao, K., Han, J., Liu, J., & Ding, E. (2021). StrucTexT: Structured text understanding with multi-modal transformers. Proceedings of the 30th International Joint Conference on Artificial Intelligence, 1383–1389.
18. Masry, A., & Hajian, A. (2024). LongFin: A multimodal document understanding model for long financial domain documents. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 1–14.
19. Mathew, M., Karatzas, D., & Jawahar, C. V. (2021). DocVQA: A dataset for VQA on document images. Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2200–2209.
20. Challa, S. R., Kaulwar, P. K., Koppolu, H. K. R., Adusupalli, B., & Suura, S. R. (2025, April). Big Data and AI in Wealth Management: Leveraging Cloud Computing for Secure and Scalable Financial Infrastructure. In International Conference on Smart Computing and Informatics (pp. 307-318). Cham: Springer Nature Switzerland.
21. Pfitzmann, B., Auer, C., Dolfi, M., Nassar, A. S., & Staar, P. (2022). DocLayNet: A large human-annotated dataset for document-layout segmentation. Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, 3743–3751.
22. Tang, Z., Yang, Z., Wang, G., Fang, Y., Liu, Y., Zhu, C., Zeng, M., Zhang, C., & Bansal, M. (2023). Unifying vision, text, and layout for universal document processing. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 19254–19264.
23. Davuluri, P. S. L. (2023). AI-Augmented Sanctions Screening: Enhancing Accuracy and Latency in Real Time Compliance Systems. AI-Augmented Sanctions Screening: Enhancing Accuracy and Latency in Real Time Compliance Systems (December 15, 2023).
24. Xu, Y., Xu, Y., Lv, T., Cui, L., Wei, F., Wang, G., Lu, Y., Florencio, D., Zhang, C., Che, W., Zhang, M., & Zhou, L. (2021). LayoutLMv2: Multi-modal pre-training for visually-rich document understanding. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, 2579–2591.
25. Zhu, F., Lei, W., Huang, Y., Wang, C., Zhang, S., Lv, J., Feng, F., & Chua, T.-S. (2021). TAT-QA: A question answering benchmark on a hybrid of tabular and textual content in finance. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, 3277–3287.
26. Vadisetty, R., Nuka, S. T., Kalisetty, S., Pandugula, C., Burugulla, J. K. R., & Annapareddy, V. N. (2025, February). Generative AI for Advanced Recycling Processes in Polyethylene and Polypropylene Manufacturing. In International Ethical Hacking Conference (pp. 269-284). Singapore: Springer Nature Singapore.
27. Xu, R., Baracaldo, N., & Joshi, J. (2021). Privacy-preserving machine learning: Methods, challenges and directions. arXiv preprint arXiv:2108.04417.
28. Singireddy, J., Kaulwar, P. K., Somu, B., Meda, R., Dodda, A., & Yellanki, S. K. (2025, August). Reinforcement Learning-Based Asset Allocation in Algorithmic Trading for Banking Institutions. In International Conference on Artificial Intelligence: Theory and Applications (pp. 357-369). Cham: Springer Nature Switzerland.
29. Li, T., Sahu, A. K., Talwalkar, A., & Smith, V. (2020). Federated learning: Challenges, methods, and future directions. IEEE Signal Processing Magazine, 37(3), 50–60.
30. Dutta, P., Adusupalli, B., Koppolu, H. K. R., Dodda, A., Yağanoğlu, M., Banerjee, J. S., & Chakraborty, A. (2025, March). Textual Social Data Disinformation Analysis Using a Hybrid Context-Enhanced Deep Learning Model. In Doctoral Symposium on Human Centered Computing (pp. 342-352). Singapore: Springer Nature Singapore.
31. Gill, S. S., Golec, M., Hu, J., Xu, M., Du, J., Wu, H., Walia, G. K., Murugesan, S. S., Ali, B., Kumar, M., Ye, K., Verma, P., Kumar, S., Cuadrado, F., & Uhlig, S. (2024). Edge AI: A taxonomy, systematic review and future directions. Internet of Things and Cyber-Physical Systems, 4, 1–25.
32. Cao, K., Chen, X., & Chen, J. (2023). Edge AI: A survey. Internet of Things and Cyber-Physical Systems, 3, 71–92.
33. Lee, Y., Park, S., & Kim, J. (2023). Deep learning-based document image analysis for intelligent document processing. Expert Systems with Applications, 213, 118958.
34. KOTHAPALLI, S. L., SEENU, A., DILEEP, V., & YASMEEN, Z. (2025). THE FUTURE OF CUSTOMER ENGAGEMENT IN RETAIL BANKING: EXPLORING THE POTENTIAL OF AUGMENTED REALITY AND IMMERSIVE TECHNOLOGIES. INTERNATIONAL JOURNAL, 73(1), 72-79.
35. Zhang, X., Wei, Y., Yang, Q., & Li, J. (2021). Document image understanding with multimodal transformer architectures. Pattern Recognition, 119, 108087.
36. Wang, X., Yu, F., Wang, B., & Zhang, Y. (2022). Transformer-based document information extraction: A survey. Information Processing & Management, 59(6), 103092.
37. Ande, R., Mehta, D., Pandugula, C., Krishna AzithTejaGanti, V., & Kalisetty, S. (2025). Modeling the Som Kamla Amba sub-watershed using Artificial Neural Networks (ANN) and the SWAT tool. Available at SSRN 5341755.
38. Li, Y., Huang, Y., Lv, T., Cui, L., & Wei, F. (2022). Layout-aware document understanding with multimodal pre-trained transformers. IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(12), 9234–9248.
39. Singh, A., Ehtesham, A., & Kumar, S. (2024). Intelligent document processing using deep learning and transformer-based architectures. Journal of King Saud University—Computer and Information Sciences, 36(2), 101–114.
40. Maguluri, K. K. (2025). Ethical challenges in artificial intelligence. How Artificial Intelligence is Transforming Healthcare IT: Applications in Diagnostics, Treatment Planning, and Patient Monitoring, 132.
41. Toprak, A., & Turan, M. (2024). Transformer-based approach for automatic semantic financial document verification. IEEE Access, 12, 184327–184349.
42. Serbanescu, V.-A., & Dhali, M. (2025). Deep learning for effective classification and information extraction of financial documents. Proceedings of the 2025 International Conference on Computer Science and Information Technology, 1–8.
43. Wang, W., Hu, H., Zhang, Z., Li, Z., Shao, H., & Dahlmeier, D. (2025). Document intelligence in the era of large language models: A survey. arXiv preprint arXiv:2510.13366.
44. Pandugula, C., Ganti, V. K. A. T., Kalisetty, S., Ande, R., & Mehta, D. (2025). Modeling the Som Kamla Amba Sub-Watershed Using Artificial Neural Networks (ANN) and the SWAT Tool. Available at SSRN 5341614.
45. Abdi, M. J., & Baghshah, M. S. (2023). Efficient document understanding with vision-language transformers. Neural Processing Letters, 55, 10457–10478.
46. Li, C., Li, R., Zhang, S., & Wang, X. (2023). Multimodal transformer networks for financial document understanding. IEEE Transactions on Artificial Intelligence, 4(6), 1548–1561.
47. Zhang, Y., Liu, Y., & Jin, L. (2021). Document layout analysis: A comprehensive survey. ACM Computing Surveys, 54(8), 1–36.
48. Xu, Y., Lv, T., Cui, L., Wang, G., Lu, Y., Florencio, D., Zhang, C., & Wei, F. (2022). XFUND: A benchmark dataset for multilingual visually rich form understanding. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics, 207–218.
49. ADUSUPALLI, B. (2025). Redefining Financial Risk Strategies: The Integration of Smart Automation, Secure Access Systems, and Predictive Intelligence in Insurance, Lending, and Asset Management. JAIBDD.
50. Lee, J., Yoon, W., Kim, S., Kim, D., Kim, S., So, C. H., & Kang, J. (2020). BioBERT: A pre-trained biomedical language representation model for biomedical text mining. Bioinformatics, 36(4), 1234–1240.
51. Howard, J., & Ruder, S. (2020). Universal language model fine-tuning for text classification. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 1–12.
52. Kairouz, P., McMahan, H. B., Avent, B., Bellet, A., Bennis, M., Bhagoji, A. N., Bonawitz, K., Charles, Z., Cormode, G., Cummings, R., D’Oliveira, R. G. L., Eichner, H., El Rouayheb, S., Evans, D., Gardner, M., Garrett, Z., Gascón, A., Ghazi, B., Gibbons, P. B., … Zhao, S. (2021). Advances and open problems in federated learning. Foundations and Trends in Machine Learning, 14(1–2), 1–210.
53. Murshed, M. G., Murphy, C., Hou, D., Khan, N., Ananthanarayanan, G., & Hussain, F. (2020). Machine learning at the network edge: A survey. ACM Computing Surveys, 53(5), 1–37.
54. Zhou, Z., Chen, X., Li, E., Zeng, L., Luo, K., & Zhang, J. (2020). Edge intelligence: Paving the last mile of artificial intelligence with edge computing. Proceedings of the IEEE, 107(8), 1738–1762.
55. Abadi, M., Chu, A., Goodfellow, I., McMahan, H. B., Mironov, I., Talwar, K., & Zhang, L. (2020). Deep learning with differential privacy: Privacy-preserving model training and inference. ACM Transactions on Privacy and Security, 23(4), 1–36.
56. Zhu, L., Liu, Z., & Han, S. (2020). Deep leakage from gradients. Advances in Neural Information Processing Systems, 33, 14774–14784.
57. Geiping, J., Bauermeister, H., Dröge, H., & Moeller, M. (2020). Inverting gradients—How easy is it to break privacy in federated learning? Advances in Neural Information Processing Systems, 33, 16937–16947.
58. Bonawitz, K., Eichner, H., Grieskamp, W., Huba, D., Ingerman, A., Ivanov, V., Kiddon, C., Konečný, J., Mazzocchi, S., McMahan, H. B., Van Overveldt, T., & Roselander, J. (2020). Towards federated learning at scale: System design. Proceedings of Machine Learning and Systems, 2, 374–388.
59. Zhou, Y., Yu, Y., Wang, Z., & Li, H. (2024). Efficient and secure edge intelligence: A survey of model compression, privacy, and deployment. IEEE Internet of Things Journal, 11(8), 13452–13471.
60. Vishnubhatla, S. (2025). Hybrid intelligence for information management systems: Converging edge AI and cloud for real-time document understanding. SSRN Electronic Journal.