Modular Data Pipeline Architecture for Large-Scale Supply Chain Optimization: A Framework for Scalable, Configuration-Driven Input Generation

Authors

  • Uday Dhembare Data Engineering Manager, Supply Chain Analytics Bellevue, WA, USA. Author

DOI:

https://doi.org/10.63282/3050-9416.IJAIBDCMS-V7I3P107

Keywords:

Data Engineering, Big Data, Automated Systems, ETL, Scalable Systems, Data Infrastructure, Supply Chain Optimization, Data Pipeline Architecture, Modular Systems, Configuration Management, Directed Acyclic Graph, Cloud-Native Computing, Network Design, Scenario Analysis, Distributed Computing, Large Scale Analytics

Abstract

Large-scale supply chain optimization relies on the steady production of accurate, validated input datasets. These datasets feed the mathematical models that handle network design, capacity planning, and cost minimization. In most organizations, the data pipelines behind these models started life as monolithic scripts—single codebases that bundle business logic, configuration parameters, and execution dependencies all in one place. That works fine at small scale, but as the number of scenarios, data sources, and downstream consumers grows, these architectures turn brittle and expensive to maintain. This paper lays out a generalized framework for rebuilding supply chain input generation pipelines around modular, configuration-driven, cloud-native principles. We draw on existing work in data engineering, distributed systems, and operations research to describe the structural problems with monolithic pipeline design, spell out the requirements for a scalable replacement, and propose an architecture built on directed acyclic graph (DAG) execution, externalized configuration management, isolated multi-scenario runs, and formal version control. We evaluate the framework against quantitative benchmarks—end-to-end runtime, scenario throughput, parameter visibility, failure recovery—and show that modular pipeline architectures cut operational overhead, improve data quality, and let organizations scale their optimization capabilities across business units and planning cycles. The implications reach across manufacturing, retail, healthcare, and logistics wherever optimization models depend on complex, multi-source input generation.

References

1. M. S. Daskin, Network and Discrete Location: Models, Algorithms, and Applications, 2nd ed. Hoboken, NJ: Wiley, 2013.

2. U. Dhembare, "Data engineering: The critical foundation for scalable supply chain optimization models," International Journal of Science and Management Studies, vol. 8, no. 2, pp. 45–62, 2025.

3. N. R. Sanders and M. Swink, "Integrating big data across the supply chain," Supply Chain Management: An International Journal, vol. 24, no. 4, pp. 1–18, 2019.

4. M. Kleppmann, Designing Data-Intensive Applications: The Big Ideas Behind Reliable, Scalable, and Maintainable Systems. Sebastopol, CA: O'Reilly Media, 2017.

5. M. Zaharia, R. S. Xin, P. Wendell, T. Das, M. Armbrust, A. Dave, X. Meng, J. Rosen, S. Venkataraman, M. J. Franklin, A. Ghodsi, J. Gonzalez, S. Shenker, and I. Stoica, "Apache Spark: A unified engine for big data processing," Communications of the ACM, vol. 59, no. 11, pp. 56–65, Nov. 2016.

6. S. Newman, Building Microservices: Designing Fine-Grained Systems, 2nd ed. Sebastopol, CA: O'Reilly Media, 2021.

7. U. Dhembare, "A progressive framework for supply chain optimization: From rule-based logic to advanced mathematical models," Kevin Open Science Journal of AIML, Data Science, Robotics, vol. 3, no. 1, pp. 12–29, 2025.

8. M. Armbrust, A. Fox, R. Griffith, A. D. Joseph, R. Katz, A. Konwinski, G. Lee, D. Patterson, A. Rabkin, I. Stoica, and M. Zaharia, "A view of cloud computing," Communications of the ACM, vol. 53, no. 4, pp. 50–58, Apr. 2010.

9. G. Wang, A. Gunasekaran, E. W. T. Ngai, and T. Papadopoulos, "Big data analytics in logistics and supply chain management: Certain investigations for research and applications," International Journal of Production Economics, vol. 176, pp. 98–110, Jun. 2016.

10. S. Tiwari, H. M. Wee, and Y. Daryanto, "Big data analytics in supply chain management between 2010 and 2016: Insights to industries," Computers & Industrial Engineering, vol. 115, pp. 319–330, Jan. 2018

Downloads

Published

2026-07-11

Issue

Section

Articles

How to Cite

1.
Dhembare U. Modular Data Pipeline Architecture for Large-Scale Supply Chain Optimization: A Framework for Scalable, Configuration-Driven Input Generation. IJAIBDCMS [Internet]. 2026 Jul. 11 [cited 2026 Aug. 20];7(3):59-65. Available from: https://ijaibdcms.org/index.php/ijaibdcms/article/view/640