A Comparative Evaluation of Learned Cardinality Estimators for Analytical Query Workloads: Accuracy, Latency, and Robustness Trade-offs
DOI:
https://doi.org/10.71222/beb4jr95Keywords:
cardinality estimation, query optimization, analytical workloads, benchmark re-analysisAbstract
Learned cardinality estimation has emerged as a promising alternative to the histogram-based methods embedded in modern query optimizers. Across the last five years, a wide range of query-driven, data-driven, and hybrid estimators have been proposed, each reporting sizeable accuracy improvements on specific benchmarks. The conditions under which these improvements translate into faster query execution, and the cost that they impose on planning and maintenance, remain less well characterized. This paper presents a cross-cutting comparative reanalysis of representative cardinality estimators, drawing performance numbers from the primary literature and four multi-method benchmark studies conducted under comparable PostgreSQL-based evaluation settings. Eight estimators are compared on accuracy; six of the eight also appear in the planning-time and training-cost comparison, for which uniformly measured numbers are available. FACE and FactorJoin are discussed as qualitative reference points, but are not in the quantitative tables. The quantitative analysis centers on JOB-light and STATS-CEB, for which comparable multi-method numbers are publicly available, with the Join Order Benchmark and the TPC-DS / DSB family providing complementary context. We compare estimators along three axes: estimation accuracy, measured by q-error percentiles; planning time and training overhead; and end-to-end query execution time on PostgreSQL with injected cardinalities. Our reanalysis confirms that data-driven estimators achieve the lowest q-error on static workloads but incur planning-time overheads that erode a portion of their end-to-end gains, and that no single estimator dominates across accuracy, latency, and resilience to data updates. Building on this analysis, we propose a workload-driven decision framework that combines a three-axis workload characterisation along query complexity, data scale, and distribution stability, a region-based selection rule, and a break-even cost-benefit model that determines when learned estimators justify their training, planning, and maintenance overhead. We discuss practical implications for analytical data pipelines and identify workload drift and planning-time efficiency as the most pressing open problems.References
1. V. Leis, A. Gubichev, A. Mirchev, P. A. Boncz, A. Kemper, and T. Neumann, "How good are query optimizers, really?," Proc. VLDB Endow., vol. 9, no. 3, pp. 204–215, 2015.
2. V. Poosala, Y. E. Ioannidis, P. J. Haas, and E. J. Shekita, "Improved histograms for selectivity estimation of range predicates," in Proc. 1996 ACM SIGMOD Int. Conf. Manage. Data, 1996, pp. 294–305.
3. A. Kipf, T. Kipf, B. Radke, V. Leis, P. Boncz, and A. Kemper, "Learned cardinalities: Estimating correlated joins with deep learning," in 9th Biennial Conf. Innovative Data Syst. Res. (CIDR), 2019.
4. Z. Yang, E. Liang, A. Kamsetty, C. Wu, Y. Duan, X. Chen, P. Abbeel, J. M. Hellerstein, S. Krishnan, and I. Stoica, "Deep unsupervised cardinality estimation," Proc. VLDB Endow., vol. 13, no. 3, pp. 279–292, 2019.
5. Z. Wu, P. Negi, M. Alizadeh, T. Kraska, and S. Madden, "FactorJoin: A new cardinality estimation framework for join queries," Proc. ACM Manag. Data, vol. 1, no. 1, pp. 41:1–41:27, 2023.
6. V. Leis, B. Radke, A. Gubichev, A. Kemper, and T. Neumann, "Cardinality estimation done right: Index-based join sampling," in 8th Biennial Conf. Innovative Data Syst. Res. (CIDR), 2017.
7. B. Hilprecht, A. Schmidt, M. Kulessa, A. Molina, K. Kersting, and C. Binnig, "DeepDB: Learn from data, not from queries!," Proc. VLDB Endow., vol. 13, no. 7, pp. 992–1005, 2020.
8. R. Zhu, Z. Wu, Y. Han, K. Zeng, A. Pfadler, Z. Qian, J. Zhou, and B. Cui, "FLAT: Fast, lightweight and accurate method for cardinality estimation," Proc. VLDB Endow., vol. 14, no. 9, pp. 1489–1502, 2021.
9. Z. Yang, A. Kamsetty, S. Luan, E. Liang, Y. Duan, X. Chen, and I. Stoica, "NeuroCard: One cardinality estimator for all tables," Proc. VLDB Endow., vol. 14, no. 1, pp. 61–73, 2021.
10. J. Wang, C. Chai, J. Liu, and G. Li, "FACE: A normalizing flow based cardinality estimator," Proc. VLDB Endow., vol. 15, no. 1, pp. 72–84, 2022.
11. P. Wu and G. Cong, "A unified deep model of learning from both data and queries for cardinality estimation," in Proc. 2021 Int. Conf. Manage. Data (SIGMOD), 2021, pp. 2009–2022.
12. K. Zhao, J. X. Yu, Z. He, R. Li, and H. Zhang, "Lightweight and accurate cardinality estimation by neural network Gaussian process," in Proc. 2022 Int. Conf. Manage. Data (SIGMOD), 2022, pp. 973–987.
13. X. Wang, C. Qu, W. Wu, J. Wang, and Q. Zhou, "Are we ready for learned cardinality estimation?," Proc. VLDB Endow., vol. 14, no. 9, pp. 1640–1654, 2021.
14. Y. Han, Z. Wu, P. Wu, R. Zhu, J. Yang, L. W. Tan, K. Zeng, G. Cong, Y. Qin, A. Pfadler, Z. Qian, J. Zhou, J. Li, and B. Cui, "Cardinality estimation in DBMS: A comprehensive benchmark evaluation," Proc. VLDB Endow., vol. 15, no. 4, pp. 752–765, 2022.
15. K. Kim, J. Jung, I. Seo, W.-S. Han, K. Choi, and J. Chong, "Learned cardinality estimation: An in-depth study," in Proc. 2022 Int. Conf. Manage. Data (SIGMOD), 2022, pp. 1214–1227.
16. J. Sun, J. Zhang, Z. Sun, G. Li, and N. Tang, "Learned cardinality estimation: A design space exploration and a comparative evaluation," Proc. VLDB Endow., vol. 15, no. 1, pp. 85–97, 2022.
17. B. Ding, S. Chaudhuri, J. Gehrke, and V. Narasayya, "DSB: A decision support benchmark for workload-driven and traditional database systems," Proc. VLDB Endow., vol. 14, no. 13, pp. 3376–3388, 2021.
18. P. Negi, Z. Wu, A. Kipf, N. Tatbul, R. Marcus, S. Madden, T. Kraska, and M. Alizadeh, "Robust query driven cardinality estimation under changing workloads," Proc. VLDB Endow., vol. 16, no. 6, pp. 1520–1533, 2023.
19. B. Li, Y. Lu, and S. Kandula, "Warper: Efficiently adapting learned cardinality estimators to data and workload drifts," in Proc. 2022 Int. Conf. Manage. Data (SIGMOD), 2022, pp. 1920–1933.
20. P. Li, W. Wei, R. Zhu, B. Ding, J. Zhou, and H. Lu, "ALECE: An attention-based learned cardinality estimator for SPJ queries on dynamic workloads," Proc. VLDB Endow., vol. 17, no. 2, pp. 197–210, 2023.
21. K. Lee, A. Dutt, V. R. Narasayya, and S. Chaudhuri, "Analyzing the impact of cardinality estimation on execution plans in Microsoft SQL Server," Proc. VLDB Endow., vol. 16, no. 11, pp. 2871–2883, 2023.
22. R. Marcus, P. Negi, H. Mao, N. Tatbul, M. Alizadeh, and T. Kraska, "Bao: Making learned query optimization practical," in Proc. 2021 Int. Conf. Manage. Data (SIGMOD), 2021, pp. 1275–1288.
23. Y. Han, H. Wang, L. Chen, Y. Dong, X. Chen, B. Yu, C. Yang, and W. Qian, "ByteCard: Enhancing ByteDance's data warehouse with learned cardinality estimation," in Companion Proc. 2024 Int. Conf. Manage. Data (SIGMOD), 2024, pp. 41–54.

