Comparative Empirical Evaluation of Prompting Strategies for Large-Language-Model Web-Navigation Agents on the WebArena Benchmark

Authors

  • Xinyu Zhang Applied Computing, University of Toronto, Toronto, ON, Canada Author
  • Yixin Zheng Information Science, University of Michigan, Ann Arbor, MI, USA Author
  • Chenhui Hao Civil Engineering, University of California, CA, USA Author

DOI:

https://doi.org/10.71222/jbq63281

Keywords:

web-navigation agents, large language models, prompting strategies, empirical evaluation

Abstract

Large language models (LLMs) are increasingly deployed as autonomous agents that operate web interfaces on behalf of users, yet reported success rates vary widely across studies that differ simultaneously in backbone model, prompting strategy, and evaluation environment, making it difficult to attribute performance to any single factor. This paper presents a controlled factorial comparison that isolates two of these factors on a fixed environment. We evaluate three backbone LLMs (GPT-4, Llama-3-70B-Instruct, and GPT-3.5) crossed with four prompting strategies (direct prompting, ReAct-style interleaved reasoning, Reflexion-style episodic retry, and best-first lookahead) on all 812 tasks of the WebArena benchmark, with a 165-instance WorkArena-L1 robustness check, under an identical action space, observation format, and step budget. Backbone choice accounts for the dominant share of variation: the weakest GPT-4 configuration (15.2%) exceeds the strongest GPT-3.5 configuration (9.7%) by 5.5 percentage points. Strategy gains are moderate and compound with backbone strength, ranging from 3.6 points on GPT-3.5 to 8.4 points on GPT-4. A cost analysis shows that the marginal cost of one additional percentage point rises roughly tenfold along the GPT-4 strategy ladder. All configurations remain far below the 78.24% human reference, indicating substantial remaining headroom.

References

1. S. Zhou, F. F. Xu, H. Zhu, X. Zhou, R. Lo, A. Sridhar, X. Cheng, T. Ou, Y. Bisk, D. Fried, U. Alon, and G. Neubig, "WebArena: A realistic web environment for building autonomous agents," in Proc. 12th Int. Conf. Learn. Representations (ICLR 2024), 2024.

2. X. Deng, Y. Gu, B. Zheng, S. Chen, S. Stevens, B. Wang, H. Sun, and Y. Su, "Mind2Web: Towards a generalist agent for the web," in Adv. Neural Inf. Process. Syst. 36 (NeurIPS 2023), Datasets and Benchmarks Track, 2023.

3. H. He, W. Yao, K. Ma, W. Yu, Y. Dai, H. Zhang, Z. Lan, and D. Yu, "WebVoyager: Building an end-to-end web agent with large multimodal models," in Proc. 62nd Annu. Meeting Assoc. Comput. Linguistics (ACL 2024), pp. 6864–6890, 2024.

4. M. Yu and Z. Li, "An Empirical Comparison of Discrete Video Tokenization Schemes for Video Question Answering and Video Captioning," Artif. Intell. Mach. Learn. Rev., vol. 6, no. 2, pp. 27–50, 2025.

5. J. Lai and Z. Li, "A Comparative Empirical Evaluation of Single-Agent and Multi-Agent LLM Prompting Strategies for Automated Formative Feedback in Education," J. Intell. Eng. Technol., vol. 1, no. 2, pp. 29–38, 2026.

6. X. Fu, T. Tang, and C. Luo, "An Empirical Comparison of ReAct, Reflexion, Plan-and-Solve, and Tree-of-Thought Planning Strategies on Financial Question Answering and Numerical Reasoning Tasks," J. Sci., Innov. Soc. Impact, vol. 2, no. 3, pp. 23–34, 2026.

7. C. Luo and X. Wang, "Latency-Throughput Tradeoffs of ONNX Runtime, TensorRT-LLM, vLLM, and Triton: An Empirical Comparison on 1B–3B Parameter LLM Inference," Spectrum Res., vol. 6, no. 1, 2026.

8. M. Li and S. Xu, "Trustworthy artificial intelligence in financial decision-making: A systematic review of explainability, fairness, and accountability," J. Sustain. Policy Pract., vol. 2, no. 3, pp. 104–114, 2026.

9. Z. Luo, "Fairness-Aware Credit Evaluation: Bias Detection and Mitigation Techniques for Inclusive Lending Practices," J. Sustain. Policy Pract., vol. 2, no. 2, pp. 78–89, 2026.

10. O. Yoran, S. J. Amouyal, C. Malaviya, B. Bogin, O. Press, and J. Berant, "AssistantBench: Can web agents solve realistic and time-consuming tasks?," in Proc. 2024 Conf. Empir. Methods Nat. Lang. Process. (EMNLP 2024), pp. 8938–8968, 2024.

11. E. Z. Liu, K. Guu, P. Pasupat, T. Shi, and P. Liang, "Reinforcement learning on web interfaces using workflow-guided exploration," in Proc. 6th Int. Conf. Learn. Representations (ICLR 2018), 2018.

12. S. Yao, H. Chen, J. Yang, and K. Narasimhan, "WebShop: Towards scalable real-world web interaction with grounded language agents," in Adv. Neural Inf. Process. Syst. 35 (NeurIPS 2022), pp. 20744–20757, 2022.

13. T. Tang and M. Yu, "A Comparative Empirical Study of Semantic Signal Enhancement Methods for User Interest Features in CTR Prediction: Applicability of TF-IDF Weighting, Sentence-BERT Embeddings, and LDA Topic Fusion," J. Comput. Innov. Appl., vol. 2, no. 1, pp. 165–174, 2024.

14. T. Tang and M. Yu, "A Comparative Evaluation of LLM-Generated Semantic Tags versus Classical Text Features (TF-IDF, LDA, BERT Embeddings) for User-Interest Enrichment in Short-Video Recommendation," Artif. Intell. Mach. Learn. Rev., vol. 5, no. 1, pp. 129–140, 2024.

15. J. Y. Koh, R. Lo, L. Jang, V. Duvvur, M. C. Lim, P.-Y. Huang, G. Neubig, S. Zhou, R. Salakhutdinov, and D. Fried, "VisualWebArena: Evaluating multimodal agents on realistic visual web tasks," in Proc. 62nd Annu. Meeting Assoc. Comput. Linguistics (ACL 2024), pp. 881–905, 2024.

16. Z. Jiang and M. Wang, "Evaluation and Analysis of Chart Reasoning Accuracy in Multimodal Large Language Models: An Empirical Study on Influencing Factors," in Pinnacle Acad. Press Proc. Ser., vol. 3, pp. 43–58, 2025.

17. Y. Tian, "Gradient Boosting-Based Demand Variability Estimation for Improved Safety Stock Calculation in Multi-Echelon Aerospace Spare Parts Inventory," J. Sci., Innov. Soc. Impact, vol. 2, no. 2, pp. 175–186, 2026.

18. X. Long, J. Hu, and Z. Ling, "A Comparative Analysis of Telemetry-Driven Anomaly Detection Approaches for Dual-Purpose Operational and Security Optimization in Edge Computing Infrastructure," J. Comput. Innov. Appl., vol. 4, no. 1, pp. 79–88, 2026.

19. A. Drouin, M. Gasse, M. Caccia, I. H. Laradji, M. Del Verme, T. Marty, L. Boisvert, M. Thakkar, Q. Cappart, D. Vazquez, N. Chapados, and A. Lacoste, "WorkArena: How capable are web agents at solving common knowledge work tasks?," in Proc. 41st Int. Conf. Mach. Learn. (ICML 2024), 2024.

20. Y. Chen and J. Hu, "Graph Neural Network-Based Cascading Disruption Path Identification in Multi-Tier Rare Earth Processing Networks," J. Global Eng. Rev., vol. 4, no. 1, pp. 99–112, 2026.

21. L. Boisvert, M. Thakkar, M. Gasse, M. Caccia, T. Le Sellier de Chezelles, Q. Cappart, N. Chapados, A. Lacoste, and A. Drouin, "WorkArena++: Towards compositional planning and reasoning-based common knowledge work tasks," in Adv. Neural Inf. Process. Syst. 37 (NeurIPS 2024), Datasets and Benchmarks Track, 2024.

22. L. Zheng, R. Wang, X. Wang, and B. An, "Synapse: Trajectory-as-exemplar prompting with memory for computer control," in Proc. 12th Int. Conf. Learn. Representations (ICLR 2024), 2024.

23. M. Wang, P. Xiao, and M. Yu, "Sparse, Dense, or Hybrid? Comparing Retrieval Strategies for Biomedical Question Answering with Retrieval-Augmented Generation," J. Sci., Innov. Soc. Impact, vol. 2, no. 2, pp. 141–152, 2026.

24. B. Zheng, B. Gou, J. Kil, H. Sun, and Y. Su, "GPT-4V(ision) is a generalist web agent, if grounded," in Proc. 41st Int. Conf. Mach. Learn. (ICML 2024), 2024.

25. Y. Chen and T. Tang, "Evaluating Prompt Engineering Strategies for Few-Shot Cyber Threat Intelligence Entity and Relation Extraction from Multi-Source Reports," J. Sci., Innov. Soc. Impact, vol. 2, no. 2, pp. 153–164, 2026.

26. J. Lai and Z. Li, "Does an LLM-Based Feedback Agent Move the Needle? An Empirical Study of Student Correctness and Hint Dependency on the ASSISTments 2017 Interaction Logs," Acad. Nexus J., vol. 5, no. 1, 2026.

27. Z. Chen and M. Wang, "Evaluating the Quality of Large Language Model-Generated Explanations in Recommendation Tasks: A Multi-Dimensional Comparative Analysis," J. Global Eng. Rev., vol. 3, no. 1, pp. 54–63, 2025.

28. T. Tang, X. Fu, and C. Luo, "An Empirical Comparison of High-Order Feature Interaction Operators for Conversion Rate Prediction in Sparse, High-Cardinality Message-Ads Traffic: Accuracy, Efficiency, and Offline-Online Consistency," J. Sci., Innov. Soc. Impact, vol. 2, no. 3, pp. 12–22, 2026.

29. X. Liu, H. Yu, H. Zhang, Y. Xu, X. Lei, H. Lai, Y. Gu, H. Ding, K. Men, K. Yang, S. Zhang, X. Deng, A. Zeng, Z. Du, C. Zhang, S. Shen, T. Zhang, Y. Su, H. Sun, M. Huang, Y. Dong, and J. Tang, "AgentBench: Evaluating LLMs as agents," in Proc. 12th Int. Conf. Learn. Representations (ICLR 2024), 2024.

30. J. Hu, X. Wang, and J. Lai, "Benchmarking Learned Cardinality Estimation Techniques for Analytical Query Processing in Data Warehouses," J. Comput. Technol. Appl. Math., vol. 3, no. 3, pp. 1–8, 2026.

31. M. Wang and L. Zhu, "Linguistic analysis of verb tense usage patterns in computer science paper abstracts," Acad. Nexus J., vol. 3, no. 3, 2024.

32. M. Wang, Z. Jiang, and S. Zhou, "Cross-cultural semantic differences in emoji usage on social media platforms," J. Adv. Comput. Syst., vol. 4, no. 5, pp. 55–66, 2024.

33. T. Le Sellier de Chezelles, M. Gasse, A. Drouin, M. Caccia, L. Boisvert, M. Thakkar, T. Marty, R. Assouel, S. Omidi Shayegan, L. K. Jang, X. H. Lu, O. Yoran, D. Kong, F. F. Xu, S. Reddy, Q. Cappart, G. Neubig, R. Salakhutdinov, N. Chapados, and A. Lacoste, "The BrowserGym ecosystem for web agent research," in Proc. Conf. Lang. Model. (COLM 2025), 2025.

34. H. Cao, J. Hu, and C. Luo, "Behavioural Feature Analysis for Anomalous Click Detection in Mobile Advertising Environments: Toward In-App Browser-Specific Detection," J. Comput. Innov. Appl., vol. 3, no. 2, pp. 96–105, 2025.

35. Y. Xu, H. Su, C. Xing, B. Mi, Q. Liu, W. Shi, B. Hui, F. Zhou, Y. Liu, T. Xie, Z. Cheng, S. Zhao, L. Kong, B. Wang, C. Xiong, and T. Yu, "Lemur: Harmonizing natural language and code for language agents," in Proc. 12th Int. Conf. Learn. Representations (ICLR 2024), 2024.

36. S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. Narasimhan, and Y. Cao, "ReAct: Synergizing reasoning and acting in language models," in Proc. 11th Int. Conf. Learn. Representations (ICLR 2023), 2023.

37. N. Shinn, F. Cassano, E. Berman, A. Gopinath, K. Narasimhan, and S. Yao, "Reflexion: Language agents with verbal reinforcement learning," in Adv. Neural Inf. Process. Syst. 36 (NeurIPS 2023), 2023.

38. J. Y. Koh, S. McAleer, D. Fried, and R. Salakhutdinov, "Tree search for language model agents," arXiv preprint arXiv:2407.01476, 2024.

39. X. H. Lu, A. Kazemnejad, N. Meade, A. Patel, D. Shin, A. Zambrano, K. Stańczak, P. Shaw, C. J. Pal, and S. Reddy, "AgentRewardBench: Evaluating automatic evaluations of web agent trajectories," arXiv preprint arXiv:2504.08942, 2025.

40. H. Lai, X. Liu, I. L. Iong, S. Yao, Y. Chen, P. Shen, H. Yu, H. Zhang, X. Zhang, Y. Dong, and J. Tang, "AutoWebGLM: A large language model-based web navigating agent," in Proc. 30th ACM SIGKDD Conf. Knowl. Discov. Data Min. (KDD 2024), pp. 5295–5306, 2024.

41. P. Sodhi, S. R. K. Branavan, Y. Artzi, and R. McDonald, "SteP: Stacked LLM policies for web actions," in Proc. Conf. Lang. Model. (COLM 2024), 2024.

42. Z. Chen, M. White, R. Mooney, A. Payani, Y. Su, and H. Sun, "When is tree search useful for LLM planning? It depends on the discriminator," in Proc. 62nd Annu. Meeting Assoc. Comput. Linguistics (ACL 2024), pp. 13659–13678, 2024.

43. F. Zhao and T. Tang, "AI-Based Sentiment Analysis for Stock Market Prediction: A Systematic Literature Review," J. Sustain. Policy Pract., vol. 2, no. 3, pp. 115–124, 2026.

44. L. Long and J. Hu, "Multi-Objective Particle Swarm Optimization for Site Selection and Policy Subsidy Maximization of Foreign Renewable Energy Enterprises in the United States," Artif. Intell. Mach. Learn. Rev., vol. 7, no. 2, pp. 54–69, 2026.

45. J. Hu and X. Long, "Graph Learning-Based Behavioral Detection for Software Supply Chain Attacks," J. Adv. Comput. Syst., vol. 4, no. 4, pp. 49–60, 2024.

46. Y. Li, F. Zhao, and J. Hu, "Identifying Cross-Market Risk Contagion Amplifiers via Graph Attention Networks: Empirical Evidence from US Financial Stress Periods," J. Comput. Innov. Appl., vol. 4, no. 1, pp. 164–175, 2026.

Downloads

Published

2026-07-03

How to Cite

Comparative Empirical Evaluation of Prompting Strategies for Large-Language-Model Web-Navigation Agents on the WebArena Benchmark. (2026). Journal of Science, Innovation & Social Impact, 2(4), 25-36. https://doi.org/10.71222/jbq63281