An Empirical Comparison of Multi-Agent LLM Collaboration Strategies for Cross-Document Fraud Evidence Aggregation and Auditable Investigation Chains
DOI:
https://doi.org/10.71222/6zavqh79Keywords:
multi-agent reasoning, fraud investigation, evidence aggregation, auditable AIAbstract
Suspicious-activity reporting under the Bank Secrecy Act requires investigators to aggregate evidence scattered across transaction records, customer identity files, behavioural logs, and external risk intelligence, while keeping every conclusion traceable to its source. Large language model agents have been proposed to support such cross-document investigation, yet little is known about how alternative collaboration strategies behave on this task. This work presents an empirical comparison of four established strategies --- single-agent serial reasoning, single-agent reasoning with self-consistency, role-specialised multi-agent collaboration, and debate-based multi-agent reasoning --- rather than proposing a new architecture. Using investigation cases synthesised from the AMLworld, Elliptic++, and Bank Account Fraud datasets, we evaluate each strategy along four axes: evidence recall, factual consistency, hallucination rate, and audit traceability. Across 300 cases on AMLworld, role-specialised collaboration attains the highest audit traceability (0.838), debate-based reasoning attains the highest factual consistency (0.841) and the lowest hallucination rate (0.094) at 6.7 times the cost of a single agent, and no strategy dominates on every axis. The gains are moderate and accompanied by substantial cost differences, indicating that strategy choice should be matched to the evidentiary and budgetary constraints of a given investigation.References
1. S. Bai, B. Wu, Y. Zhang, C. Wu, X. Zheng, Y. Yuan, K. Wu, and J. Li, "AuditAgent: Expert-guided multi-agent reasoning for cross-document fraudulent evidence discovery," in Proc. 6th ACM Int. Conf. AI in Finance (ICAIF '25). New York, NY, USA: Association for Computing Machinery, 2025. [Online]. Available: https://arxiv.org/abs/2510.00156
2. L. Huang, W. Yu, W. Ma, W. Zhong, Z. Feng, H. Wang, Q. Chen, W. Peng, X. Feng, B. Qin, and T. Liu, "A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions," ACM Trans. Inf. Syst., vol. 43, no. 2, Art. no. 42, 2025. doi: 10.1145/3703155
3. Y. Du, S. Li, A. Torralba, J. B. Tenenbaum, and I. Mordatch, "Improving factuality and reasoning in language models through multiagent debate," in Proc. 41st Int. Conf. Mach. Learn. (ICML 2024). Proceedings of Machine Learning Research, 2024. [Online]. Available: https://arxiv.org/abs/2305.14325
4. Q. Xie et al., "FinBen: A holistic financial benchmark for large language models," in Adv. Neural Inf. Process. Syst. 37 (NeurIPS 2024) Datasets and Benchmarks Track, 2024. [Online]. Available: https://arxiv.org/abs/2402.12659
5. Q. Wu, G. Bansal, J. Zhang, Y. Wu, S. Zhang, E. Zhu, B. Li, L. Jiang, X. Zhang, and C. Wang, "AutoGen: Enabling next-gen LLM applications via multi-agent conversation," arXiv preprint arXiv:2308.08155, 2023. [Online]. Available: https://arxiv.org/abs/2308.08155
6. S. Hong, M. Zhuge, J. Chen, X. Zheng, Y. Cheng, J. Wang, C. Zhang, Z. Wang, S. K. S. Yau, Z. Lin, L. Zhou, C. Ran, L. Xiao, C. Wu, and J. Schmidhuber, "MetaGPT: Meta programming for a multi-agent collaborative framework," in Proc. 12th Int. Conf. Learn. Represent. (ICLR 2024), 2024. [Online]. Available: https://arxiv.org/abs/2308.00352
7. T. Liang, Z. He, W. Jiao, X. Wang, Y. Wang, R. Wang, Y. Yang, S. Shi, and Z. Tu, "Encouraging divergent thinking in large language models through multi-agent debate," in Proc. 2024 Conf. Empirical Methods Nat. Lang. Process. (EMNLP 2024). Stroudsburg, PA, USA: Association for Computational Linguistics, 2024, pp. 17889–17904. [Online]. Available: https://arxiv.org/abs/2305.19118
8. X. Wang, J. Wei, D. Schuurmans, Q. Le, E. Chi, S. Narang, A. Chowdhery, and D. Zhou, "Self-consistency improves chain of thought reasoning in language models," in Proc. 11th Int. Conf. Learn. Represent. (ICLR 2023), 2023. [Online]. Available: https://arxiv.org/abs/2203.11171
9. X. Liu et al., "AgentBench: Evaluating LLMs as agents," in Proc. 12th Int. Conf. Learn. Represent. (ICLR 2024), 2024. [Online]. Available: https://arxiv.org/abs/2308.03688
10. S. Min, K. Krishna, X. Lyu, M. Lewis, W. Yih, P. W. Koh, M. Iyyer, L. Zettlemoyer, and H. Hajishirzi, "FActScore: Fine-grained atomic evaluation of factual precision in long form text generation," in Proc. 2023 Conf. Empirical Methods Nat. Lang. Process. (EMNLP 2023). Stroudsburg, PA, USA: Association for Computational Linguistics, 2023, pp. 12076–12100. [Online]. Available: https://arxiv.org/abs/2305.14251
11. M. Weber, G. Domeniconi, J. Chen, D. K. I. Weidele, C. Bellei, T. Robinson, and C. E. Leiserson, "Anti-money laundering in Bitcoin: Experimenting with graph convolutional networks for financial forensics," in KDD '19 Workshop Anomaly Detect. Finance, 2019. [Online]. Available: https://arxiv.org/abs/1908.02591
12. B. Egressy, L. von Niederhäusern, J. Blanuša, E. Altman, R. Wattenhofer, and K. Atasu, "Provably powerful graph neural networks for directed multigraphs," in Proc. 38th AAAI Conf. Artif. Intell. (AAAI-24), vol. 38, no. 10, 2024, pp. 11838–11846. doi: 10.1609/aaai.v38i10.29069
13. P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. Küttler, M. Lewis, W. Yih, T. Rocktäschel, S. Riedel, and D. Kiela, "Retrieval-augmented generation for knowledge-intensive NLP tasks," in Adv. Neural Inf. Process. Syst. 33 (NeurIPS 2020), 2020, pp. 9459–9474. [Online]. Available: https://arxiv.org/abs/2005.11401
14. E. Altman, J. Blanuša, L. von Niederhäusern, B. Egressy, A. Anghel, and K. Atasu, "Realistic synthetic financial transactions for anti-money laundering models," in Adv. Neural Inf. Process. Syst. 36 (NeurIPS 2023) Datasets and Benchmarks Track, 2023. [Online]. Available: https://arxiv.org/abs/2306.16424
15. Y. Elmougy and L. Liu, "Demystifying fraudulent transactions and illicit nodes in the Bitcoin network for financial forensics," in Proc. 29th ACM SIGKDD Conf. Knowl. Discov. Data Mining (KDD '23). New York, NY, USA: Association for Computing Machinery, 2023. doi: 10.1145/3580305.3599803
16. G. Li, H. A. A. K. Hammoud, H. Itani, D. Khizbullin, and B. Ghanem, "CAMEL: Communicative agents for 'mind' exploration of large language model society," in Adv. Neural Inf. Process. Syst. 36 (NeurIPS 2023), 2023, pp. 51991–52008. [Online]. Available: https://arxiv.org/abs/2303.17760
17. C. Qian, W. Liu, H. Liu, N. Chen, Y. Dang, J. Li, C. Yang, W. Chen, Y. Su, X. Cong, J. Xu, D. Li, Z. Liu, and M. Sun, "ChatDev: Communicative agents for software development," in Proc. 62nd Annu. Meeting Assoc. Comput. Linguistics (ACL 2024). Stroudsburg, PA, USA: Association for Computational Linguistics, 2024, pp. 15174–15186. [Online]. Available: https://arxiv.org/abs/2307.07924
18. C.-M. Chan, W. Chen, Y. Su, J. Yu, W. Xue, S. Zhang, J. Fu, and Z. Liu, "ChatEval: Towards better LLM-based evaluators through multi-agent debate," in Proc. 12th Int. Conf. Learn. Represent. (ICLR 2024), 2024. [Online]. Available: https://arxiv.org/abs/2308.07201
19. L. Zheng et al., "Judging LLM-as-a-judge with MT-Bench and Chatbot Arena," in Adv. Neural Inf. Process. Syst. 36 (NeurIPS 2023) Datasets and Benchmarks Track, 2023. [Online]. Available: https://arxiv.org/abs/2306.05685
20. Y. Dou, Z. Liu, L. Sun, Y. Deng, H. Peng, and P. S. Yu, "Enhancing graph neural network-based fraud detectors against camouflaged fraudsters," in Proc. 29th ACM Int. Conf. Inf. Knowl. Manage. (CIKM '20). New York, NY, USA: Association for Computing Machinery, 2020, pp. 315–324. doi: 10.1145/3340531.3411903

