Distributed Index Optimization Techniques for Time-Sensitive Massive Data
DOI:
https://doi.org/10.71222/4xthr117Keywords:
time-series data, distributed indexing, data sharding, hierarchical storage, query optimizationAbstract
With the continuous deployment of industrial perception networks and smart urban monitoring systems, the volume of temporal stream data has expanded significantly, evolving from terabyte-level to petabyte-scale volumes. The high-throughput and persistent nature of data generation imposes stringent demands on storage and retrieval architectures. Traditional single-machine indexing architectures suffer from limited computational power and storage capacity, making them inadequate for processing ultra-large-scale temporal datasets. Existing distributed indexing solutions predominantly employ fixed sharding mechanisms, often encountering challenges such as uneven node load distribution, high query interaction costs across nodes, and inefficient reuse of cold and hot data indices, thereby failing to simultaneously ensure stable high-concurrency writes and low-latency multi-dimensional queries. Addressing these technical challenges, this paper proposes a dual-layer hybrid distributed indexing architecture that integrates temporal data characteristics with hierarchical storage techniques. The proposed approach compresses index metadata using temporal feature hash summaries, implements dynamic scheduling and partition pruning mechanisms based on data access frequency, and optimizes the underlying logic for log-structured merge tree merging operations. Experimental results across multiple benchmark datasets confirm the effectiveness of the proposed solution in reducing indexing overhead, substantially lowering cross-partition query latency, and achieving balanced cluster workload distribution. The findings provide a viable and scalable technical framework for the efficient distributed retrieval of massive temporal datasets in real-world industrial and urban monitoring applications.References
1. L. Zhang, N. Alghamdi, M. Y. Eltabakh, and E. A. Rundensteiner, "TARDIS: Distributed indexing framework for big time series data," in 2019 IEEE 35th International Conference on Data Engineering (ICDE), Apr. 2019, pp. 1202–1213.
2. D. E. Yagoubi, R. Akbarinia, F. Masseglia, and T. Palpanas, "Massively distributed time series indexing and querying," IEEE Trans. Knowl. Data Eng., vol. 32, no. 1, pp. 108–120, 2018.
3. N. Alghamdi, L. Zhang, H. Zhang, E. A. Rundensteiner, and M. Y. Eltabakh, "Chainlink: indexing big time series data for long subsequence matching," in 2020 IEEE 36th International Conference on Data Engineering (ICDE), Apr. 2020, pp. 529–540.
4. B. Greenstein, S. Ratnasamy, S. Shenker, R. Govindan, and D. Estrin, "DIFS: A distributed index for features in sensor networks," Ad Hoc Netw., vol. 1, no. 2–3, pp. 333–349, 2003.
5. M. Malensek, S. Pallickara, and S. Pallickara, "Analytic queries over geospatial time-series data using distributed hash tables," IEEE Trans. Knowl. Data Eng., vol. 28, no. 6, pp. 1408–1422, 2016.
6. D. E. Yagoubi, R. Akbarinia, F. Masseglia, and D. Shasha, "Radiussketch: massively distributed indexing of time series," in 2017 IEEE International Conference on Data Science and Advanced Analytics (DSAA), Oct. 2017, pp. 262–271.
7. M. Hojati, S. Roberts, and C. Robertson, "Dstree: A spatio-temporal indexing data structure for distributed networks," Math. Comput. Appl., vol. 29, no. 3, p. 42, 2024.
8. O. Levchenko, D. E. Yagoubi, R. Akbarinia, F. Masseglia, B. Kolev, and D. Shasha, "Spark-parsketch: a massively distributed indexing of time series datasets," in Proc. 27th ACM Int. Conf. Inf. Knowl. Manag., Oct. 2018, pp. 1951–1954.
9. X. Huang, J. Wang, R. Wong, J. Zhang, and C. Wang, "Pisa: An index for aggregating big time series data," in Proc. 25th ACM Int. Conf. Inf. Knowl. Manag., Oct. 2016, pp. 979–988.
10. M. Younan, E. H. Houssein, M. Elhoseny, and A. E. M. Ali, "Performance analysis for similarity data fusion model for enabling time series indexing in internet of things applications," PeerJ Comput. Sci., vol. 7, p. e500, 2021.
11. T. Palpanas, "The parallel and distributed future of data series mining," in 2017 International Conference on High Performance Computing & Simulation (HPCS), Jul. 2017, pp. 916–920.
12. X. Liu, Z. Wei, W. Yu, S. Liu, G. Wang, X. Liu, and Y. Li, "Khronos: A real-time indexing framework for time series databases on large-scale performance monitoring systems," in Proc. 32nd ACM Int. Conf. Inf. Knowl. Manag., Oct. 2023, pp. 1607–1616.
13. G. Chatzigeorgakidis, D. Skoutas, K. Patroumpas, S. Athanasiou, and S. Skiadopoulos, "Indexing geolocated time series data," in Proc. 25th ACM SIGSPATIAL Int. Conf. Adv. Geogr. Inf. Syst., Nov. 2017, pp. 1–10.
14. C. W. Tan, G. I. Webb, and F. Petitjean, "Indexing and classifying gigabytes of time series under time warping," in Proc. 2017 SIAM Int. Conf. Data Min., Jun. 2017, pp. 282–290.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Ruohua Yuan (Author)

This work is licensed under a Creative Commons Attribution 4.0 International License.

