Latency-Aware LLM Architectures for Financial Decision Systems

Authors

  • Michael Anderson Author

Keywords:

Generative Intelligence, Enterprise, Finance, Foundation Models, Prompt Engineering, Backward Process, Oversight, Data Hastings, Pipeline Construction, Cost Modeling, Game Theory.

Abstract

Enterprise financial decision-making increasingly depends on systems that are both intelligent and performance-aware. The rise of decision pipelines within enterprise finance has created new opportunities in what can be called the Enterprise Generative Decision-AI space. Generative Artificial Intelligence (GAI) has attracted significant attention for its success in image synthesis, text generation, code generation, and related tasks. Where discriminative models estimate the conditional probability of an outcome given a set of features, GAI models instead learn the full joint distribution of the data, leveraging the conditional independence structure inherent in probability distributions. This distinction gives GAI strong potential across a range of enterprise financial applications, including forecasting, business strategy formulation, risk assessment, decision-space exploration in complex environments, and broader decision-support functions.

Realizing this potential at scale, however, is complicated by the nature of modern financial data. Enterprises must contend with time-series data, unstructured sources such as text, and structured data drawn from trading, treasury, and risk management systems — all of which must be processed quickly and cost-effectively to support low-latency decision-making. These demands make a monolithic approach to enterprise GAI difficult to sustain, and instead call for system-wide architectural patterns designed for scale. In particular, latency-sensitive workloads require Systems-Aware Generative AI Models that balance responsiveness against compute cost. One such pattern, Containerized Financial AI-as-a-Service, aggregates GAI and other model endpoints across the enterprise cloud, enabling financial prediction and decision-exploration workloads to run with low latency and transparent, well-defined resource costing.

References

1. Chen, Z., Chen, W., Smiley, C., Shah, S., Borova, I., Langdon, D., Moussa, R., Beane, M., Huang, T.-H., Routledge, B., & Wang, W. Y. (2021). FinQA: A dataset of numerical reasoning over financial data. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (pp. 3697–3711). Association for Computational Linguistics.

2. Zhu, F., Lei, W., Huang, Y., Wang, C., Zhang, S., Lv, J., Feng, F., & Chua, T.-S. (2021). TAT-QA: A question answering benchmark on a hybrid of tabular and textual content in finance. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers) (pp. 3277–3287). Association for Computational Linguistics.

3. Nandan, B. P., & Chitta, S. S. (2023). Machine Learning Driven Metrology and Defect Detection in Extreme Ultraviolet (EUV) Lithography: A Paradigm Shift in Semiconductor Manufacturing. Educational Administration: Theory and Practice, 29 (4), 4555–4568. International Journal of Scientific Research and Modern Technology, 1(12), 216-226.

4. So, D., Mańke, W., Liu, H., Dai, Z., Shazeer, N., & Le, Q. V. (2021). Searching for efficient transformers for language modeling. In Advances in Neural Information Processing Systems, 34. Neural Information Processing Systems Foundation.

5. Chuang, C.-Y., & Yang, Y. (2022). Buy Tesla, sell Ford: Assessing implicit stock market preference in pre-trained language models. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) (pp. 100–105). Association for Computational Linguistics.

6. Shah, R., Chawla, K., Eidnani, D., Shah, A., Du, W., Chava, S., Raman, N., Smiley, C., Chen, J., & Yang, D. (2022). When FLUE meets FLANG: Benchmarks and large pretrained language model for financial domain. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing (pp. 10359–10371). Association for Computational Linguistics.

7. Chen, Z., Li, S., Smiley, C., Ma, Z., Shah, S., & Wang, W. Y. (2022). ConvFinQA: Exploring the chain of numerical reasoning in conversational finance question answering. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing (pp. 6279–6292). Association for Computational Linguistics.

8. Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., & Chen, W. (2022). LoRA: Low-rank adaptation of large language models. In International Conference on Learning Representations.

9. Dao, T., Fu, D. Y., Ermon, S., Rudra, A., & Ré, C. (2022). FlashAttention: Fast and memory-efficient exact attention with IO-awareness. In Advances in Neural Information Processing Systems, 35. Neural Information Processing Systems Foundation.

10. Frantar, E., Ashkboos, S., Hoefler, T., & Alistarh, D. (2022). GPTQ: Accurate post-training quantization for generative pre-trained transformers. arXiv preprint arXiv:2210.17323.

11. Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., de Las Casas, D., Hendricks, L. A., Welbl, J., Clark, A., Hennigan, T., Noland, E., Millican, K., van den Driessche, G., Damoc, B., Guy-Brown, C., Osindero, S., Simonyan, K., Elsen, E., ... Sifre, L. (2022). Training compute-optimal large language models. In Advances in Neural Information Processing Systems, 35. Neural Information Processing Systems Foundation.

12. Adusupalli, B., & Insurity-Lead, A. C. E. (2024). The role of internal audit in enhancing corporate governance: A comparative analysis of risk management and compliance strategies. Outcomes. Journal for ReAttach Therapy and Developmental Diversities, 6, 1921-1937.

13. Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., Schuh, P., Shi, K., Tsvyashchenko, S., Maynez, J., Rao, A., Barnes, P., Tay, Y., Shazeer, N., Prabhakaran, V., ... Fiedel, N. (2022). PaLM: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311.

14. Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P., Leike, J., & Lowe, R. (2022). Training language models to follow instructions with human feedback. In Advances in Neural Information Processing Systems, 35. Neural Information Processing Systems Foundation.

15. Xiao, G., Lin, J., Seznec, M., Wu, H., Demouth, J., & Han, S. (2023). SmoothQuant: Accurate and efficient post-training quantization for large language models. In Proceedings of the 40th International Conference on Machine Learning (pp. 38087–38099). PMLR.

16. Wu, S., Irsoy, O., Lu, S., Dabravolski, V., Dredze, M., Gehrmann, S., Kambadur, P., Rosenberg, D., & Mann, G. (2023). BloombergGPT: A large language model for finance. arXiv preprint arXiv:2303.17564.

17. Yang, H., Liu, X.-Y., & Wang, C. D. (2023). FinGPT: Open-source financial large language models. arXiv preprint arXiv:2306.06031.

18. Li, X., Chan, S., Zhu, X., Pei, Y., Ma, Z., Liu, X., & Shah, S. (2023). Are ChatGPT and GPT-4 general-purpose solvers for financial text analytics? A study on several typical tasks. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing: Industry Track (pp. 3924–3937). Association for Computational Linguistics.

19. Guo, Y., Xu, Z., & Yang, Y. (2023). Is ChatGPT a financial expert? Evaluating language models on financial natural language processing. In Findings of the Association for Computational Linguistics: EMNLP 2023 (pp. 815–821). Association for Computational Linguistics.

20. Mashetty, S. (2024). The role of US patents and trademarks in advancing mortgage financing technologies. European Advanced Journal for Science & Engineering (EAJSE)-p-ISSN, 3050-9696.

21. Rodriguez Inserte, P., Nakhlé, M., Qader, R., Caillaut, G., & Liu, J. (2023). Large language model adaptation for financial sentiment analysis. In Proceedings of the Sixth Workshop on Financial Technology and Natural Language Processing (pp. 1–10). Association for Computational Linguistics.

22. Rajpoot, P., & Parikh, A. (2023). GPT-FinRE: In-context learning for financial relation extraction using large language models. In Proceedings of the Sixth Workshop on Financial Technology and Natural Language Processing (pp. 42–45). Association for Computational Linguistics.

23. Li, Y., Wang, S., Ding, H., & Chen, H. (2023). Large language models in finance: A survey. arXiv preprint arXiv:2311.10723.

24. Lee, J., Stevens, N., Han, S. C., & Song, M. (2024). A survey of large language models in finance (FinLLMs). arXiv preprint arXiv:2402.02315.

25. Bhatia, G., Nagoudi, E. M. B., Cavusoglu, H., & Abdul-Mageed, M. (2024). FinTral: A family of GPT-4 level multimodal financial large language models. In Findings of the Association for Computational Linguistics: ACL 2024 (pp. 13064–13087). Association for Computational Linguistics.

26. Kirtac, K., & Germano, G. (2024). Enhanced financial sentiment analysis and trading strategy development using large language models. In Proceedings of the 14th Workshop on Computational Approaches to Subjectivity, Sentiment, & Social Media Analysis (pp. 1–10). Association for Computational Linguistics.

27. Aguda, T. D., Siddagangappa, S., Kochkina, E., Kaur, S., Wang, D., & Smiley, C. (2024). Large language models as financial data annotators: A study on effectiveness and efficiency. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) (pp. 10124–10145). ELRA and ICCL.

28. Recharla, M., & Chitta, S. AI-Enhanced Neuroimaging and Deep Learning-Based Early Diagnosis of Multiple Sclerosis and Alzheimer’s.

29. Xie, Q., Han, W., Chen, Z., Xiang, R., Zhang, X., He, Y., Xiao, M., Li, D., Dai, Y., Feng, D., Xu, Y., Kang, H., Kuang, Z., Yuan, C., Yang, K., Luo, Z., Zhang, T., Liu, Z., Xiong, G., ... Huang, J. (2024). FinBen: A holistic financial benchmark for large language models. In Advances in Neural Information Processing Systems, 37. Neural Information Processing Systems Foundation.

30. Xie, Q., Huang, J., Li, D., Chen, Z., Xiang, R., Xiao, M., Yu, Y., Somasundaram, V., Yang, K., Yuan, C., Luo, Z., Liu, Z., He, Y., Jiang, Y., Li, H., Feng, D., Liu, X.-Y., Wang, B., Lopez-Lira, A., ... Lai, Y. (2024). FinNLP-AgentScen-2024 shared task: Financial challenges in large language models—FinLLMs. In Proceedings of the Eighth Financial Technology and Natural Language Processing and the 1st Agent AI for Scenario Planning Workshop. Association for Computational Linguistics.

31. Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J., Zhang, H., & Stoica, I. (2023). Efficient memory management for large language model serving with PagedAttention. In Proceedings of the 29th Symposium on Operating Systems Principles (pp. 611–626). Association for Computing Machinery.

32. Leviathan, Y., Kalman, M., & Matias, Y. (2023). Fast inference from transformers via speculative decoding. In Proceedings of the 40th International Conference on Machine Learning (Vol. 202, pp. 19274–19286). PMLR.

33. Lin, J., Tang, J., Tang, H., Yang, S., Chen, W.-M., Wang, W.-C., Xiao, G., Dang, X., Gan, C., & Han, S. (2023). AWQ: Activation-aware weight quantization for LLM compression and acceleration. arXiv preprint arXiv:2306.00978.

34. Dettmers, T., Pagnoni, A., Holtzman, A., & Zettlemoyer, L. (2023). QLoRA: Efficient finetuning of quantized LLMs. In Advances in Neural Information Processing Systems, 36. Neural Information Processing Systems Foundation.

35. Dao, T. (2024). FlashAttention-2: Faster attention with better parallelism and work partitioning. In International Conference on Learning Representations.

36. Paleti, S. (2024). Data engineering for AI-powered compliance: A new paradigm in banking risk management. Available at SSRN 5256619.

37. Cai, T., Li, Y., Geng, Z., Peng, H., Lee, J. D., Chen, D., & Dao, T. (2024). Medusa: Simple LLM inference acceleration framework with multiple decoding heads. In Proceedings of the 41st International Conference on Machine Learning (Vol. 235, pp. 5209–5235). PMLR.

38. Zhong, Y., Liu, S., Chen, J., Hu, J., Zhu, Y., Liu, X., Jin, X., & Zhang, H. (2024). DistServe: Disaggregating prefill and decoding for goodput-optimized large language model serving. arXiv preprint arXiv:2401.09670.

39. Agrawal, A., Kedia, N., Panwar, A., Mohan, J., Kwatra, N., Gulavani, B. S., Tumanov, A., & Ramjee, R. (2024). Taming throughput-latency tradeoff in LLM inference with Sarathi-Serve. In 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24) (pp. 117–134). USENIX Association.

40. Prabhakar, R. B., Zhang, H., & Wentzlaff, D. (2024). Kraken: Inherently parallel transformers for efficient multi-device inference. In Advances in Neural Information Processing Systems, 37. Neural Information Processing Systems Foundation.

Additional Files

Published

2025-02-17

Data Availability Statement

None

How to Cite

Latency-Aware LLM Architectures for Financial Decision Systems. (2025). European Data Science Journal (EDSJ), 3(01). https://esa-research.org/index.php/EDSJ/article/view/210

Most read articles by the same author(s)

Similar Articles

41-50 of 52

You may also start an advanced similarity search for this article.