Designing Modular AI Architectures for Enterprise Scale: From Monolithic Models to Composable AI Systems

Main Article Content

Hemant Soni
Priyadharsini Ramamurthy

Abstract

Enterprises adopting machine learning have moved beyond isolated proof of concept models towards production systems that must integrate with established platforms such as customer relationship management (CRM), enterprise resource planning (ERP) and telecom operations and business support systems (OSS/BSS). This paper argues that the dominant obstacle to value realisation is architectural rather than algorithmic. We review the limitations of the monolithic model pattern, in which a single large model is trained, deployed and maintained as one indivisible unit, and we set out a reference architecture for modular, composable AI systems. The proposed architecture organises capability into independently deployable model services, a shared data foundation with a feature store and lineage tracking, and an operations control plane that automates integration, testing, monitoring and governance. We map established software engineering practice, namely microservices, continuous delivery and design patterns for machine learning, onto the constraints of enterprise AI, and we discuss integration with CRM, ERP and telecom stacks through an API and event driven serving layer. A qualitative evaluation compares the monolithic and modular approaches across maintainability, scalability, governance and time to change. The analysis indicates that modularisation reduces coupling and shortens the path from experimentation to a governed, observable production system, at the cost of an increased operational surface area that the control plane is designed to manage

Article Details

Section

Articles

How to Cite

Designing Modular AI Architectures for Enterprise Scale: From Monolithic Models to Composable AI Systems. (2023). International Journal of Research Publications in Engineering, Technology and Management (IJRPETM), 6(1), 8136-8141. https://doi.org/10.15662/IJRPETM.2023.0601013

References

[1] D. Sculley, G. Holt, D. Golovin, E. Davydov, T. Phillips, D. Ebner, V. Chaudhary, M. Young, J. F. Crespo, and D. Dennison, "Hidden Technical Debt in Machine Learning Systems," in Advances in Neural Information Processing Systems (NeurIPS), vol. 28, 2015, pp. 2503-2511.

[2] E. Breck, S. Cai, E. Nielsen, M. Salib, and D. Sculley, "The ML Test Score: A Rubric for ML Production Readiness and Technical Debt Reduction," in Proc. IEEE Int. Conf. on Big Data, 2017, pp. 1123-1132.

[3] S. Amershi, A. Begel, C. Bird, R. DeLine, H. Gall, E. Kamar, N. Nagappan, B. Nushi, and T. Zimmermann, "Software Engineering for Machine Learning: A Case Study," in Proc. IEEE/ACM 41st Int. Conf. on Software Engineering: Software Engineering in Practice (ICSE-SEIP), 2019, pp. 291-300.

[4] D. Kreuzberger, N. Kuhl, and S. Hirschl, "Machine Learning Operations (MLOps): Overview, Definition, and Architecture," arXiv preprint arXiv:2205.02302, 2022.

[5] M. M. John, H. H. Olsson, and J. Bosch, "Towards MLOps: A Framework and Maturity Model," in Proc. 47th Euromicro Conf. on Software Engineering and Advanced Applications (SEAA), 2021, pp. 1-8.

[6] D. Baylor, E. Breck, H.-T. Cheng, N. Fiedel, C. Y. Foo, Z. Haque, S. Haykal, M. Ispir, V. Jain, L. Koc, et al., "TFX: A TensorFlow-Based Production-Scale Machine Learning Platform," in Proc. 23rd ACM SIGKDD Int. Conf. on Knowledge Discovery and Data Mining (KDD), 2017, pp. 1387-1395.

[7] V. Lakshmanan, S. Robinson, and M. Munn, Machine Learning Design Patterns. Sebastopol, CA, USA: O’Reilly Media, 2020.

[8] S. Newman, Building Microservices: Designing Fine-Grained Systems. Sebastopol, CA, USA: O’Reilly Media, 2015.

[9] N. Dragoni, S. Giallorenzo, A. L. Lafuente, M. Mazzara, F. Montesi, R. Mustafin, and L. Safina, "Microservices: Yesterday, Today, and Tomorrow," in Present and Ulterior Software Engineering, Springer, 2017, pp. 195-216.

[10] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, "Attention Is All You Need," in Advances in Neural Information Processing Systems (NeurIPS), vol. 30, 2017, pp. 5998-6008.

[11] T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, et al., "Language Models are Few-Shot Learners," in Advances in Neural Information Processing Systems (NeurIPS), vol. 33, 2020, pp. 1877-1901.

[12] J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, "BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding," in Proc. NAACL-HLT, 2019, pp. 4171-4186.

[13] R. Bommasani, D. A. Hudson, E. Adeli, R. Altman, S. Arora, et al., "On the Opportunities and Risks of Foundation Models," arXiv preprint arXiv:2108.07258, 2021.

[14] P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. Kuttler, M. Lewis, W. Yih, T. Rocktaschel, S. Riedel, and D. Kiela, "Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks," in Advances in Neural Information Processing Systems (NeurIPS), vol. 33, 2020, pp. 9459-9474.

[15] A. Paleyes, R.-G. Urma, and N. D. Lawrence, "Challenges in Deploying Machine Learning: A Survey of Case Studies," arXiv preprint arXiv:2011.09926, 2020.