Quantifying the Cost and Risk of Enterprise LLM Adaptation Strategies
Main Article Content
Abstract
Enterprises choosing between prompt engineering, retrieval-augmented generation (RAG) and parameter-efficient fine-tuning are well served by the accuracy literature and poorly served by everything else. Published comparisons establish which technique injects knowledge more effectively in a given domain, but none reports what the alternatives cost to build and run, how they differ in latency, what infrastructure each obliges an organisation to own, or how their risk exposures differ in kind rather than degree. This paper supplies that missing quantitative and operational account. We first consolidate the comparative accuracy evidence and state plainly what it does and does not support. We then develop an adaptation cost of ownership model and derive a closed-form break-even condition for fine-tuning relative to prompting. The condition has three properties that are counter-intuitive and consequential: output-token cost cancels entirely, so the economic case rests only on input tokens removed per request; the break-even volume is inversely proportional to input token price, so the same task amortises very differently on frontier and small models; and it is inversely proportional to the achieved prompt-length reduction, which is a design variable rather than a constant. Under illustrative parameters the break-even sits near 0.64 million requests per month for a 4,000-token prompt, well above the volume of most internal enterprise applications. The practical conclusion is that fine-tuning is a behavioural instrument rather than a cost-reduction strategy, and that RAG lowers cost at no volume relative to the prompt it augments, so its business case must rest on freshness, provenance and accuracy. We close with a latency decomposition showing that decode time dominates, and a risk analysis showing that the archetypes do not sit on a single risk gradient: fine-tuning concentrates data risk, retrieval concentrates system risk, and prompting concentrates assurance risk.
Article Details
Section
How to Cite
References
[1] K. Nisar, “Prompting, Retrieval, and Fine-Tuning: Foundations of Enterprise Language Model Adaptation,” International Journal of Engineering & Extended Technologies Research (IJEETR), vol. 6, no. 6, pp. 9310–9319, Nov. 2024, doi: 10.15662/IJEETR.2024.0606030.
[2] O. Ovadia, M. Brief, M. Mishaeli, and O. Elisha, “Fine-tuning or retrieval? Comparing knowledge injection in LLMs,” arXiv preprint arXiv:2312.05934, Dec. 2023.
[3] A. Balaguer, V. Benara, R. L. de Freitas Cunha, R. de M. Estevão Filho, T. Hendry, D. Holstein, J. Marsman, N. Mecklenburg, S. Malvar, L. O. Nunes, R. Padilha, M. Sharp, B. Silva, S. Sharma, V. Aski, and R. Chandra, “RAG vs fine-tuning: Pipelines, tradeoffs, and a case study on agriculture,” arXiv preprint arXiv:2401.08406, Jan. 2024.
[4] H. Soudani, E. Kanoulas, and F. Hasibi, “Fine tuning vs. retrieval augmented generation for less popular knowledge,” arXiv preprint arXiv:2403.01432, Mar. 2024.
[5] K. Nisar, “Bridging Prototyping and Production: A Systems-Level Study of ML Workflow Standardization in Enterprise AI Adoption,” International Journal of Computational and Experimental Science and Engineering, vol. 11, no. 4, 2025, doi: 10.22399/ijcesen.5364.
[6] T. Zhang, S. G. Patil, N. Jain, S. Shen, M. Zaharia, I. Stoica, and J. E. Gonzalez, “RAFT: Adapting language model to domain specific RAG,” arXiv preprint arXiv:2403.10131, Mar. 2024.
[7] N. Mecklenburg, Y. Lin, X. Li, D. Holstein, L. Nunes, S. Malvar, B. Silva, R. Chandra, V. Aski, P. K. R. Yannam, T. Aktas, and T. Hendry, “Injecting new knowledge into large language models via supervised fine-tuning,” arXiv preprint arXiv:2404.00213, Mar. 2024.
[8] D. Biderman, J. Portes, J. J. Gonzalez Ortiz, M. Paul, P. Greengard, C. Jennings, D. King, S. Havens, V. Chiley, J. Frankle, C. Blakeney, and J. P. Cunningham, “LoRA learns less and forgets less,” arXiv preprint arXiv:2405.09673, May 2024.
[9] S. Es, J. James, L. Espinosa Anke, and S. Schockaert, “RAGAs: Automated evaluation of retrieval augmented generation,” in Proc. 18th Conf. European Chapter of the Association for Computational Linguistics: System Demonstrations, Mar. 2024, pp. 150–158, doi: 10.18653/v1/2024.eacl-demo.16.
[10] J. Chen, H. Lin, X. Han, and L. Sun, “Benchmarking large language models in retrieval-augmented generation,” in Proc. AAAI Conf. Artificial Intelligence, vol. 38, no. 16, 2024, pp. 17754–17762, doi: 10.1609/aaai.v38i16.29728.
[11] S. Barnett, S. Kurniawan, S. Thudumu, Z. Brannelly, and M. Abdelrazek, “Seven failure points when engineering a retrieval augmented generation system,” in Proc. IEEE/ACM 3rd Int. Conf. AI Engineering: Software Engineering for AI (CAIN), 2024, pp. 194–199, doi: 10.1145/3644815.3644945.
[12] W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. E. Gonzalez, H. Zhang, and I. Stoica, “Efficient memory management for large language model serving with PagedAttention,” in Proc. 29th ACM Symp. Operating Systems Principles (SOSP), 2023, pp. 611–626, doi: 10.1145/3600006.3613165.
[13] Y. Leviathan, M. Kalman, and Y. Matias, “Fast inference from transformers via speculative decoding,” in Proc. 40th Int. Conf. Machine Learning (ICML), PMLR vol. 202, 2023, pp. 19274–19286.
[14] Y. A. Malkov and D. A. Yashunin, “Efficient and robust approximate nearest neighbor search using hierarchical navigable small world graphs,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 42, no. 4, pp. 824–836, Apr. 2020, doi: 10.1109/TPAMI.2018.2889473.
[15] D. Sculley, G. Holt, D. Golovin, E. Davydov, T. Phillips, D. Ebner, V. Chaudhary, M. Young, J.-F. Crespo, and D. Dennison, “Hidden technical debt in machine learning systems,” in Advances in Neural Information Processing Systems 28, 2015, pp. 2503–2511.
[16] A. Paleyes, R.-G. Urma, and N. D. Lawrence, “Challenges in deploying machine learning: A survey of case studies,” ACM Computing Surveys, vol. 55, no. 6, art. 114, pp. 1–29, 2023, doi: 10.1145/3533378.
[17] T. Dettmers, A. Pagnoni, A. Holtzman, and L. Zettlemoyer, “QLoRA: Efficient finetuning of quantized LLMs,” in Advances in Neural Information Processing Systems 36, 2023, pp. 10088–10115.
[18] J. Johnson, M. Douze, and H. Jégou, “Billion-scale similarity search with GPUs,” IEEE Transactions on Big Data, vol. 7, no. 3, pp. 535–547, Jul. 2021, doi: 10.1109/TBDATA.2019.2921572.
[19] R. Nogueira and K. Cho, “Passage re-ranking with BERT,” arXiv preprint arXiv:1901.04085, Jan. 2019.
[20] N. Carlini, F. Tramèr, E. Wallace, M. Jagielski, A. Herbert-Voss, K. Lee, A. Roberts, T. Brown, D. Song, Ú. Erlingsson, A. Oprea, and C. Raffel, “Extracting training data from large language models,” in Proc. 30th USENIX Security Symposium, 2021, pp. 2633–2650.
[21] N. Carlini, D. Ippolito, M. Jagielski, K. Lee, F. Tramèr, and C. Zhang, “Quantifying memorization across neural language models,” in Proc. Int. Conf. Learning Representations (ICLR), 2023.
[22] National Institute of Standards and Technology, Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1, Gaithersburg, MD, USA, Jan. 2023, doi: 10.6028/NIST.AI.100-1.
[23] T. Gebru, J. Morgenstern, B. Vecchione, J. W. Vaughan, H. Wallach, H. Daumé III, and K. Crawford, “Datasheets for datasets,” Communications of the ACM, vol. 64, no. 12, pp. 86–92, Dec. 2021, doi: 10.1145/3458723.
[24] K. Greshake, S. Abdelnabi, S. Mishra, C. Endres, T. Holz, and M. Fritz, “Not what you’ve signed up for: Compromising real-world LLM-integrated applications with indirect prompt injection,” in Proc. 16th ACM Workshop on Artificial Intelligence and Security (AISec), 2023, pp. 79–90, doi: 10.1145/3605764.3623985.
[25] Y. Yao, J. Duan, K. Xu, Y. Cai, Z. Sun, and Y. Zhang, “A survey on large language model (LLM) security and privacy: The Good, The Bad, and The Ugly,” High-Confidence Computing, vol. 4, no. 2, art. 100211, Jun. 2024, doi: 10.1016/j.hcc.2024.100211.
[26] M. Sclar, Y. Choi, Y. Tsvetkov, and A. Suhr, “Quantifying language models’ sensitivity to spurious features in prompt design, or: How I learned to start worrying about prompt formatting,” in Proc. Int. Conf. Learning Representations (ICLR), 2024.
[27] European Parliament and Council of the European Union, Regulation (EU) 2024/1689 of 13 June 2024 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act), OJ L, 2024/1689, 12 Jul. 2024.
[28] I. M. Enholm, E. Papagiannidis, P. Mikalef, and J. Krogstie, “Artificial intelligence and business value: A literature review,” Information Systems Frontiers, vol. 24, no. 5, pp. 1709–1734, Oct. 2022, doi: 10.1007/s10796-021-10186-w.
[29] P. Mikalef and M. Gupta, “Artificial intelligence capability: Conceptualization, measurement calibration, and empirical study on its impact on organizational creativity and firm performance,” Information & Management, vol. 58, no. 3, art. 103434, Apr. 2021, doi: 10.1016/j.im.2021.103434.