{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T18:55:21Z","timestamp":1777488921053,"version":"3.51.4"},"reference-count":99,"publisher":"Association for Computing Machinery (ACM)","issue":"4","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Intell. Syst. Technol."],"published-print":{"date-parts":[[2026,8,31]]},"abstract":"<jats:p>\n                    The AI community has witnessed the emergence of various chat-style Large Language Models (LLMs) since the advent of ChatGPT. Despite significant progress in this area, evaluating these models remains a substantial challenge. The evaluations provided by humans or GPT-4 oracles are often taken as the gold standard, but they are neither automatic nor scalable. More recently, a series of (open source) LLM-based judge models have been introduced, yet they often exhibit model-specific biases, e.g., a LLaMA-family judge favors a LLaMA-family model. On the other hand, autoregressive evaluation metrics, which holds the potential to address the aforementioned issues, remains underexplored. Among them, likelihood-based metrics such as perplexity and Negative Log-Likelihood (NLL) are widely adopted and has proven effective in tracking the pre-training progress of LLMs. However, they struggle to evaluate the generation capabilities of fine-tuned models due to\n                    <jats:italic toggle=\"yes\">exposure bias<\/jats:italic>\n                    , a phenomenon where the distribution of the model\u2019s output gradually deviates from the ground-truth during inference. To address this key issue, in this article, we propose a novel autoregressive metric, Normalized Discounted Cumulative Gain (NDCG), to improve the evaluation of fine-tuned LLMs. Our experimental results demonstrate that NDCG significantly outperforms likelihood-based metrics: it shows over 45% improvement in both Spearman and Kendall\u2019s tau correlation coefficients for commonsense QA tasks, and aligns more closely with GPT-4 Elo rankings for instruction-tuned models.\n                  <\/jats:p>","DOI":"10.1145\/3763000","type":"journal-article","created":{"date-parts":[[2025,10,29]],"date-time":"2025-10-29T14:05:41Z","timestamp":1761746741000},"page":"1-27","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["An Improved Autoregressive Evaluation Paradigm for Large Language Models"],"prefix":"10.1145","volume":"17","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-3269-6992","authenticated-orcid":false,"given":"Jipeng","family":"Zhang","sequence":"first","affiliation":[{"name":"The Hong Kong University of Science and Technology, Hong Kong, Hong Kong"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7217-0656","authenticated-orcid":false,"given":"Rui","family":"Pan","sequence":"additional","affiliation":[{"name":"The Hong Kong University of Science and Technology, Hong Kong, Hong Kong"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-7313-6108","authenticated-orcid":false,"given":"Yuzheng","family":"Hu","sequence":"additional","affiliation":[{"name":"University of Illinois Urbana-Champaign, Urbana, Illinois, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-5168-8345","authenticated-orcid":false,"given":"KaShun","family":"Shum","sequence":"additional","affiliation":[{"name":"The Hong Kong University of Science and Technology, Hong Kong, Hong Kong"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-8642-3860","authenticated-orcid":false,"given":"Guanyu","family":"Yao","sequence":"additional","affiliation":[{"name":"University of California Santa Barbara, Santa Barbara, California, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0003-2639-1804","authenticated-orcid":false,"given":"Xiang","family":"Liu","sequence":"additional","affiliation":[{"name":"The Hong Kong University of Science and Technology, Hong Kong, Hong Kong"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5033-7096","authenticated-orcid":false,"given":"Renjie","family":"Pi","sequence":"additional","affiliation":[{"name":"The Hong Kong University of Science and Technology, Hong Kong, Hong Kong"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8846-1260","authenticated-orcid":false,"given":"Hanze","family":"Dong","sequence":"additional","affiliation":[{"name":"The Hong Kong University of Science and Technology, Hong Kong, Hong Kong"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3325-9209","authenticated-orcid":false,"given":"Shizhe","family":"Diao","sequence":"additional","affiliation":[{"name":"The Hong Kong University of Science and Technology, Hong Kong, Hong Kong"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3525-6738","authenticated-orcid":false,"given":"Yong","family":"Lin","sequence":"additional","affiliation":[{"name":"The Hong Kong University of Science and Technology, Hong Kong, Hong Kong"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8579-1600","authenticated-orcid":false,"given":"Han","family":"Zhao","sequence":"additional","affiliation":[{"name":"University of Illinois Urbana-Champaign, Urbana, Illinois, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5511-2558","authenticated-orcid":false,"given":"Tong","family":"Zhang","sequence":"additional","affiliation":[{"name":"University of Illinois Urbana-Champaign, Urbana, Illinois, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2026,4,28]]},"reference":[{"key":"e_1_3_1_2_2","unstructured":"Stability AI. 2023. StableLM: Stability AI Language Models. Retrieved from https:\/\/github.com\/Stability-AI\/StableLM"},{"key":"e_1_3_1_3_2","doi-asserted-by":"crossref","unstructured":"Ebtesam Almazrouei Hamza Alobeidli Abdulaziz Alshamsi Alessandro Cappelli Ruxandra Cojocaru Merouane Debbah Etienne Goffinet Daniel Heslow Julien Launay Quentin Malartic et al. 2023. The RefinedWeb dataset for Falcon LLM: outperforming curated corpora with web data and web data only. arXiv:2306.01116. Retrieved from https:\/\/arxiv.org\/abs\/2306.01116","DOI":"10.52202\/075280-3464"},{"key":"e_1_3_1_4_2","unstructured":"Rohan Anil Andrew M. Dai Orhan Firat Melvin Johnson Dmitry Lepikhin Alexandre Passos Siamak Shakeri Emanuel Taropa Paige Bailey Zhifeng Chen et al. 2023. Palm 2 technical report. arXiv:2305.10403. Retrieved from https:\/\/arxiv.org\/abs\/2305.10403"},{"key":"e_1_3_1_5_2","unstructured":"Anthropic. 2022. Introducing Claude. Retrieved from https:\/\/www.anthropic.com\/index\/introducing-claude"},{"key":"e_1_3_1_6_2","unstructured":"Zhangir Azerbayev Hailey Schoelkopf Keiran Paster Marco Dos Santos Stephen McAleer Albert Q. Jiang Jia Deng Stella Biderman and Sean Welleck. 2023. Llemma: An open language model for mathematics. arXiv:2310.10631. Retrieved from https:\/\/arxiv.org\/abs\/2310.10631"},{"key":"e_1_3_1_7_2","unstructured":"Jinze Bai Shuai Bai Yunfei Chu Zeyu Cui Kai Dang Xiaodong Deng Yang Fan Wenbin Ge Yu Han Fei Huang et al. 2023. Qwen technical report. arXiv:2309.16609. Retrieved from https:\/\/arxiv.org\/abs\/2309.16609"},{"key":"e_1_3_1_8_2","unstructured":"Yuntao Bai Andy Jones Kamal Ndousse Amanda Askell Anna Chen Nova DasSarma Dawn Drain Stanislav Fort Deep Ganguli Tom Henighan et al. 2022. Training a helpful and harmless assistant with reinforcement learning from human feedback. Retrieved from arXiv:2204.05862. Retrieved from https:\/\/arxiv.org\/abs\/2204.05862"},{"key":"e_1_3_1_9_2","unstructured":"Edward Beeching Cl\u00e9mentine Fourrier Nathan Habib Sheon Han Nathan Lambert Nazneen Rajani Omar Sanseviero Lewis Tunstall and Thomas Wolf. 2023. Open LLM Leaderboard. Retrieved from https:\/\/huggingface.co\/spaces\/HuggingFaceH4\/open_llm_leaderboard"},{"key":"e_1_3_1_10_2","unstructured":"Stella Biderman Hailey Schoelkopf Quentin Anthony Herbie Bradley Kyle O\u2019Brien Eric Hallahan Mohammad Aflah Khan Shivanshu Purohit Usvsn Sai Prashanth Edward Raff et al. 2023. Pythia: A suite for analyzing large language models across training and scaling. arXiv:2304.01373. Retrieved from https:\/\/arxiv.org\/abs\/2304.01373"},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1006\/csla.2000.0150"},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i05.6239"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","unstructured":"Sid Black Leo Gao Phil Wang Connor Leahy and Stella Biderman. 2021. GPT-Neo: Large Scale Autoregressive Language Modeling with Mesh-Tensorflow. DOI: 10.5281\/zenodo.5297715","DOI":"10.5281\/zenodo.5297715"},{"key":"e_1_3_1_14_2","first-page":"2206","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Borgeaud Sebastian","year":"2022","unstructured":"Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George Bm Van Den Driessche, Jean-Baptiste Lespiau, Bogdan Damoc, Aidan Clark, et al. 2022. Improving language models by retrieving from trillions of tokens. In Proceedings of the International Conference on Machine Learning. PMLR, 2206\u20132240."},{"key":"e_1_3_1_15_2","first-page":"1877","article-title":"Language models are few-shot learners","volume":"33","author":"Brown Tom","year":"2020","unstructured":"Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D. Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems, Vol. 33 (2020), 1877\u20131901.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_16_2","unstructured":"Stephen Casper Xander Davies Claudia Shi Thomas Krendl Gilbert J\u00e9r\u00e9my Scheurer Javier Rando Rachel Freedman Tomasz Korbak David Lindner Pedro Freire et al. 2023. Open problems and fundamental limitations of reinforcement learning from human feedback. arXiv:2307.15217. Retrieved from https:\/\/arxiv.org\/abs\/2307.15217"},{"key":"e_1_3_1_17_2","volume-title":"Proceedings of the 12th International Conference on Learning Representations","author":"Chan Chi-Min","year":"2023","unstructured":"Chi-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu, Wei Xue, Shanghang Zhang, Jie Fu, and Zhiyuan Liu. 2023. ChatEval: Towards better LLM-based evaluators through multi-agent debate. In Proceedings of the 12th International Conference on Learning Representations."},{"key":"e_1_3_1_18_2","doi-asserted-by":"crossref","unstructured":"Ciprian Chelba Tomas Mikolov Mike Schuster Qi Ge Thorsten Brants Phillipp Koehn and Tony Robinson. 2013. One billion word benchmark for measuring progress in statistical language modeling. arXiv:1312.3005. Retrieved from https:\/\/arxiv.org\/abs\/1312.3005","DOI":"10.21437\/Interspeech.2014-564"},{"key":"e_1_3_1_19_2","unstructured":"Qinyuan Cheng Tianxiang Sun Wenwei Zhang Siyin Wang Xiangyang Liu Mozhi Zhang Junliang He Mianqiu Huang Zhangyue Yin Kai Chen et al. 2023. Evaluating hallucinations in chinese large language models. arXiv:2310.03368. Retrieved from https:\/\/arxiv.org\/abs\/2310.03368"},{"key":"e_1_3_1_20_2","unstructured":"Wei Lin Chiang Zhuohan Li Zi Lin Ying Sheng Zhanghao Wu Hao Zhang Lianmin Zheng Siyuan Zhuang Yonghao Zhuang Joseph E. Gonzalez Ion Stoica and Eric P. Xing. 2023. Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90%* ChatGPT Quality. Retrieved from https:\/\/lmsys.org\/blog\/2023-03-30-vicuna\/"},{"key":"e_1_3_1_21_2","unstructured":"Christopher Clark Kenton Lee Ming-Wei Chang Tom Kwiatkowski Michael Collins and Kristina Toutanova. 2019. BoolQ: Exploring the surprising difficulty of natural yes\/no questions. arXiv:1905.10044. Retrieved from https:\/\/arxiv.org\/abs\/1905.10044"},{"key":"e_1_3_1_22_2","unstructured":"Peter Clark Isaac Cowhey Oren Etzioni Tushar Khot Ashish Sabharwal Carissa Schoenick and Oyvind Tafjord. 2018. Think you have solved question answering? Try arc the AI2 reasoning challenge. arXiv:1803.05457. Retrieved from https:\/\/arxiv.org\/abs\/1803.05457"},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1006\/csla.2000.0156"},{"key":"e_1_3_1_24_2","unstructured":"Databricks. 2023. Dolly. Retrieved from https:\/\/github.com\/databrickslabs\/dolly"},{"key":"e_1_3_1_25_2","doi-asserted-by":"crossref","unstructured":"Tim Dettmers Artidoro Pagnoni Ari Holtzman and Luke Zettlemoyer. 2023. Qlora: Efficient finetuning of quantized LLMs. arXiv:2305.14314. Retrieved from https:\/\/arxiv.org\/abs\/2305.14314","DOI":"10.52202\/075280-0441"},{"issue":"8","key":"e_1_3_1_26_2","first-page":"242","article-title":"The proposed USCF rating system. Its development, theory, and applications","volume":"22","author":"Elo Arpad E.","year":"1967","unstructured":"Arpad E. Elo. 1967. The proposed USCF rating system. Its development, theory, and applications. Chess Life 22, 8 (1967), 242\u2013247.","journal-title":"Chess Life"},{"key":"e_1_3_1_27_2","volume-title":"The Rating of Chessplayers, Past and Present","author":"Elo Arpad E.","year":"1978","unstructured":"Arpad E. Elo. 1978. The Rating of Chessplayers, Past and Present. Arco Publishing."},{"key":"e_1_3_1_28_2","unstructured":"Suriya Gunasekar Yi Zhang Jyoti Aneja Caio C\u00e9sar Teodoro Mendes Allie Del Giorno Sivakanth Gopi Mojan Javaheripi Piero Kauffmann Gustavo de Rosa Olli Saarikivi et al. 2023. Textbooks are all you need. arXiv:2306.11644. Retrieved from https:\/\/arxiv.org\/abs\/2306.11644"},{"key":"e_1_3_1_29_2","unstructured":"Zhen Guo Peiqi Wang Yanwei Wang and Shangdi Yu. 2023. LLaMA: Improving small language models in domain-specific QA via generative data augmentation. arXiv:2305.07804. Retrieved from https:\/\/arxiv.org\/abs\/2305.07804"},{"key":"e_1_3_1_30_2","first-page":"5549","volume-title":"Proceedings of the International Conference on Artificial Intelligence and Statistics","author":"Hegselmann Stefan","year":"2023","unstructured":"Stefan Hegselmann, Alejandro Buendia, Hunter Lang, Monica Agrawal, Xiaoyi Jiang, and David Sontag. 2023. TabLLM: Few-shot classification of tabular data with large language models. In Proceedings of the International Conference on Artificial Intelligence and Statistics. PMLR, 5549\u20135581."},{"key":"e_1_3_1_31_2","unstructured":"Dan Hendrycks Collin Burns Steven Basart Andy Zou Mantas Mazeika Dawn Song and Jacob Steinhardt. 2020. Measuring massive multitask language understanding. arXiv:2009.03300. Retrieved from https:\/\/arxiv.org\/abs\/2009.03300"},{"key":"e_1_3_1_32_2","unstructured":"Danny Hernandez Jared Kaplan Tom Henighan and Sam McCandlish. 2021. Scaling laws for transfer. arXiv:2102.01293. Retrieved from https:\/\/arxiv.org\/abs\/2102.01293"},{"key":"e_1_3_1_33_2","unstructured":"Jordan Hoffmann Sebastian Borgeaud Arthur Mensch Elena Buchatskaya Trevor Cai Eliza Rutherford Diego de Las Casas Lisa Anne Hendricks Johannes Welbl Aidan Clark et al. 2022. Training compute-optimal large language models. arXiv:2203.15556. Retrieved from https:\/\/arxiv.org\/abs\/2203.15556"},{"key":"e_1_3_1_34_2","volume-title":"Proceedings of the 10th International Conference on Learning Representations (ICLR \u201922)","author":"Jang Joel","year":"2022","unstructured":"Joel Jang, Seonghyeon Ye, Sohee Yang, Joongbo Shin, Janghoon Han, Gyeonghun Kim, Stanley Jungkyu Choi, and Minjoon Seo. 2022. Towards continual knowledge learning of language models. In Proceedings of the 10th International Conference on Learning Representations (ICLR \u201922)."},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.1145\/582415.582418"},{"key":"e_1_3_1_36_2","unstructured":"Albert Q. Jiang Alexandre Sablayrolles Arthur Mensch Chris Bamford Devendra Singh Chaplot Diego de las Casas Florian Bressand Gianna Lengyel Guillaume Lample Lucile Saulnier et al. 2023. Mistral 7B. arXiv:2310.06825. Retrieved from https:\/\/arxiv.org\/abs\/2310.06825"},{"key":"e_1_3_1_37_2","unstructured":"Albert Q. Jiang Alexandre Sablayrolles Antoine Roux Arthur Mensch Blanche Savary Chris Bamford Devendra Singh Chaplot Diego de las Casas Emma Bou Hanna Florian Bressand et al. 2024. Mixtral of experts. arXiv:2401.04088. Retrieved from https:\/\/arxiv.org\/abs\/2401.04088"},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P17-1147"},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00276"},{"key":"e_1_3_1_40_2","doi-asserted-by":"crossref","unstructured":"Yanis Labrak Adrien Bazoge Emmanuel Morin Pierre-Antoine Gourraud Mickael Rouvier and Richard Dufour. 2024. BioMistral: A collection of open-source pretrained large language models for medical domains. arXiv:2402.10373. Retrieved from https:\/\/arxiv.org\/abs\/2402.10373","DOI":"10.18653\/v1\/2024.findings-acl.348"},{"key":"e_1_3_1_41_2","unstructured":"Dawei Li Bohan Jiang Liangjie Huang Alimohammad Beigi Chengshuai Zhao Zhen Tan Amrita Bhattacharjee Yuxuan Jiang Canyu Chen Tianhao Wu et al. 2024. From generation to judgment: Opportunities and challenges of LLM-as-a-judge. arXiv:2411.16594. Retrieved from https:\/\/arxiv.org\/abs\/2411.16594"},{"key":"e_1_3_1_42_2","unstructured":"Junlong Li Shichao Sun Weizhe Yuan Run-Ze Fan Hai Zhao and Pengfei Liu. 2023. Generative judge for evaluating alignment. arXiv:2310.05470. Retrieved from https:\/\/arxiv.org\/abs\/2310.05470"},{"key":"e_1_3_1_43_2","unstructured":"Raymond Li Loubna Ben Allal Yangtian Zi Niklas Muennighoff Denis Kocetkov Chenghao Mou Marc Marone Christopher Akiki Jia Li Jenny Chim et al. 2023. StarCoder: May the source be with you! arXiv:2305.06161. Retrieved from https:\/\/arxiv.org\/abs\/2305.06161"},{"key":"e_1_3_1_44_2","unstructured":"Xuechen Li Tianyi Zhang Yann Dubois Rohan Taori Ishaan Gulrajani Carlos Guestrin Percy Liang and Tatsunori B. Hashimoto. 2023. AlpacaEval: An Automatic Evaluator of Instruction-Following Models. Retrieved from https:\/\/github.com\/tatsu-lab\/alpaca_eval"},{"key":"e_1_3_1_45_2","unstructured":"Yuanzhi Li S\u00e9bastien Bubeck Ronen Eldan Allie Del Giorno Suriya Gunasekar and Yin Tat Lee. 2023. Textbooks are all you need II: phi-1.5 technical report. arXiv:2309.05463. Retrieved from https:\/\/arxiv.org\/abs\/2309.05463"},{"key":"e_1_3_1_46_2","unstructured":"Percy Liang Rishi Bommasani Tony Lee Dimitris Tsipras Dilara Soylu Michihiro Yasunaga Yian Zhang Deepak Narayanan Yuhuai Wu Ananya Kumar et al. 2022. Holistic evaluation of language models. arXiv:2211.09110. Retrieved from https:\/\/arxiv.org\/abs\/2211.09110"},{"key":"e_1_3_1_47_2","first-page":"74","volume-title":"Text Summarization Branches Out","author":"Lin Chin-Yew","year":"2004","unstructured":"Chin-Yew Lin. 2004. ROUGE: A package for automatic evaluation of summaries. In Text Summarization Branches Out. Association for Computational Linguistics, Barcelona, Spain, 74\u201381."},{"key":"e_1_3_1_48_2","unstructured":"Haotian Liu Chunyuan Li Qingyang Wu and Yong Jae Lee. 2023. Visual instruction tuning. arXiv:2304.08485. Retrieved from https:\/\/arxiv.org\/abs\/2304.08485"},{"key":"e_1_3_1_49_2","first-page":"22188","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Liu Hong","year":"2023","unstructured":"Hong Liu, Sang Michael Xie, Zhiyuan Li, and Tengyu Ma. 2023. Same pre-training loss, better downstream: Implicit bias matters for language models. In Proceedings of the International Conference on Machine Learning. PMLR, 22188\u201322214."},{"key":"e_1_3_1_50_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICDM.2011.118"},{"key":"e_1_3_1_51_2","unstructured":"Ian Magnusson Akshita Bhagia Valentin Hofmann Luca Soldaini Ananya Harsh Jha Oyvind Tafjord Dustin Schwenk Evan Pete Walsh Yanai Elazar Kyle Lo et al. 2023. Paloma: A benchmark for evaluating language model fit. arXiv:2312.10523. Retrieved from https:\/\/arxiv.org\/abs\/2312.10523"},{"key":"e_1_3_1_52_2","doi-asserted-by":"publisher","DOI":"10.5555\/972470.972475"},{"key":"e_1_3_1_53_2","doi-asserted-by":"crossref","unstructured":"Todor Mihaylov Peter Clark Tushar Khot and Ashish Sabharwal. 2018. Can a suit of armor conduct electricity? A new dataset for open book question answering. arXiv:1809.02789. Retrieved from https:\/\/arxiv.org\/abs\/1809.02789","DOI":"10.18653\/v1\/D18-1260"},{"key":"e_1_3_1_54_2","unstructured":"Behrad Moniri Hamed Hassani and Edgar Dobriban. 2024. Evaluating the performance of large language models via debates. arXiv:2406.11044. Retrieved from https:\/\/arxiv.org\/abs\/2406.11044"},{"key":"e_1_3_1_55_2","unstructured":"Erik Nijkamp Bo Pang Hiroaki Hayashi Lifu Tu Huan Wang Yingbo Zhou Silvio Savarese and Caiming Xiong. 2022. CodeGen: An open large language model for code with multi-turn program synthesis. arXiv:2203.13474. Retrieved from https:\/\/arxiv.org\/abs\/2203.13474"},{"key":"e_1_3_1_56_2","unstructured":"OpenAI. 2022. Introducing ChatGPT. Retrieved from https:\/\/openai.com\/blog\/chatgpt"},{"key":"e_1_3_1_57_2","unstructured":"OpenAI. 2023. GPT-4 technical report. arXiv:2303.08774. Retrieved from https:\/\/arxiv.org\/abs\/2303.08774"},{"key":"e_1_3_1_58_2","doi-asserted-by":"publisher","DOI":"10.3115\/1073083.1073135"},{"key":"e_1_3_1_59_2","doi-asserted-by":"publisher","DOI":"10.3115\/1073083.1073135"},{"key":"e_1_3_1_60_2","unstructured":"Guilherme Penedo Quentin Malartic Daniel Hesslow Ruxandra Cojocaru Alessandro Cappelli Hamza Alobeidli Baptiste Pannier Ebtesam Almazrouei and Julien Launay. 2023. The RefinedWeb dataset for Falcon LLM: Outperforming curated corpora with web data and web data only. arXiv:2306.01116. Retrieved from https:\/\/arxiv.org\/abs\/2306.01116"},{"key":"e_1_3_1_61_2","unstructured":"Baolin Peng Chunyuan Li Pengcheng He Michel Galley and Jianfeng Gao. 2023. Instruction tuning with GPT-4. arXiv:2304.03277. Retrieved from https:\/\/arxiv.org\/abs\/2304.03277"},{"key":"e_1_3_1_62_2","unstructured":"Renjie Pi Jiahui Gao Shizhe Diao Rui Pan Hanze Dong Jipeng Zhang Lewei Yao Jianhua Han Hang Xu Lingpeng Kong et al. 2023. DetGPT: Detect what you need via reasoning. arXiv:2305.14167. Retrieved from https:\/\/arxiv.org\/abs\/2305.14167"},{"key":"e_1_3_1_63_2","unstructured":"Jack W. Rae Sebastian Borgeaud Trevor Cai Katie Millican Jordan Hoffmann Francis Song John Aslanides Sarah Henderson Roman Ring Susannah Young et al. 2021. Scaling language models: Methods analysis & insights from training gopher. arXiv:2112.11446. Retrieved from https:\/\/arxiv.org\/abs\/2112.11446"},{"key":"e_1_3_1_64_2","volume-title":"Proceedings of the 4th International Conference on Learning Representations (ICLR \u201916)","author":"Ranzato Marc\u2019Aurelio","year":"2016","unstructured":"Marc\u2019Aurelio Ranzato, Sumit Chopra, Michael Auli, and Wojciech Zaremba. 2016. Sequence level training with recurrent neural networks. In Proceedings of the 4th International Conference on Learning Representations (ICLR \u201916)."},{"issue":"9","key":"e_1_3_1_65_2","doi-asserted-by":"crossref","first-page":"99","DOI":"10.1145\/3474381","article-title":"Winogrande: An adversarial winograd schema challenge at scale","volume":"64","author":"Sakaguchi Keisuke","year":"2021","unstructured":"Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2021. Winogrande: An adversarial winograd schema challenge at scale. Communication of the ACM 64, 9 (2021), 99\u2013106.","journal-title":"Communication of the ACM"},{"key":"e_1_3_1_66_2","unstructured":"Teven Le Scao Angela Fan Christopher Akiki Ellie Pavlick Suzana Ilic Daniel Hesslow Roman Castagn\u00e9 Alexandra Sasha Luccioni Fran\u00e7ois Yvon Matthias Gall\u00e9 et al. 2022. BLOOM: A 176B-parameter open-access multilingual language model. arXiv:2211.05100. Retrieved from https:\/\/arxiv.org\/abs\/2211.05100"},{"key":"e_1_3_1_67_2","unstructured":"Rylan Schaeffer Brando Miranda and Sanmi Koyejo. 2023. Are emergent abilities of large language models a mirage? arXiv:2304.15004. Retrieved from https:\/\/arxiv.org\/abs\/2304.15004"},{"key":"e_1_3_1_68_2","doi-asserted-by":"crossref","unstructured":"Thibault Sellam Dipanjan Das and Ankur P. Parikh. 2020. BLEURT: Learning robust metrics for text generation. arXiv:2004.04696. Retrieved from https:\/\/arxiv.org\/abs\/2004.04696","DOI":"10.18653\/v1\/2020.acl-main.704"},{"key":"e_1_3_1_69_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-0-387-85820-3_8"},{"key":"e_1_3_1_70_2","unstructured":"Aarohi Srivastava Abhinav Rastogi Abhishek Rao Abu Awal Md Shoeb Abubakar Abid Adam Fisch Adam R. Brown Adam Santoro Aditya Gupta Adri\u00e0 Garriga-Alonso et al. 2022. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. arXiv:2206.04615. Retrieved from https:\/\/arxiv.org\/abs\/2206.04615"},{"key":"e_1_3_1_71_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.emnlp-main.62"},{"key":"e_1_3_1_72_2","first-page":"8968","volume-title":"Proceedings of the 34th AAAI Conference on Artificial Intelligence (AAAI \u201920), 32nd Innovative Applications of Artificial Intelligence Conference (IAAI \u201920), 10th AAAI Symposium on Educational Advances in Artificial Intelligence (EAAI \u201920)","author":"Sun Yu","year":"2020","unstructured":"Yu Sun, Shuohuan Wang, Yu-Kun Li, Shikun Feng, Hao Tian, Hua Wu, and Haifeng Wang. 2020. ERNIE 2.0: A continual pre-training framework for language understanding. In Proceedings of the 34th AAAI Conference on Artificial Intelligence (AAAI \u201920), 32nd Innovative Applications of Artificial Intelligence Conference (IAAI \u201920), 10th AAAI Symposium on Educational Advances in Artificial Intelligence (EAAI \u201920). AAAI Press, 8968\u20138975."},{"key":"e_1_3_1_73_2","unstructured":"Ross Taylor Marcin Kardas Guillem Cucurull Thomas Scialom Anthony Hartshorn Elvis Saravia Andrew Poulton Viktor Kerkez and Robert Stojnic. 2022. Galactica: A large language model for science. arXiv:2211.09085. Retrieved from https:\/\/arxiv.org\/abs\/2211.09085"},{"key":"e_1_3_1_74_2","unstructured":"MosaicML NLP Team. 2023. Introducing MPT-7B: A New Standard for Open-Source Commercially Usable LLMs. Retrieved from https:\/\/www.mosaicml.com\/blog\/mpt-7b"},{"key":"e_1_3_1_75_2","unstructured":"TogetherAI. 2023. Releasing 3B and 7B RedPajama-INCITE Family of Models Including Base Instruction-Tuned & Chat Models. Retrieved from https:\/\/www.together.xyz\/blog\/redpajama-models-v1"},{"key":"e_1_3_1_76_2","unstructured":"Hugo Touvron Thibaut Lavril Gautier Izacard Xavier Martinet Marie-Anne Lachaux Timoth\u00e9e Lacroix Baptiste Rozi\u00e8re Naman Goyal Eric Hambro Faisal Azhar et al. 2023. LLaMA: Open and efficient foundation language models. arXiv:2302.13971. Retrieved from https:\/\/arxiv.org\/abs\/2302.13971"},{"key":"e_1_3_1_77_2","unstructured":"Hugo Touvron Louis Martin Kevin Stone Peter Albert Amjad Almahairi Yasmine Babaei Nikolay Bashlykov Soumya Batra Prajjwal Bhargava Shruti Bhosale et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv:2307.09288. Retrieved from https:\/\/arxiv.org\/abs\/2307.09288"},{"key":"e_1_3_1_78_2","doi-asserted-by":"crossref","unstructured":"Alex Wang Amanpreet Singh Julian Michael Felix Hill Omer Levy and Samuel R. Bowman. 2018. GLUE: A multi-task benchmark and analysis platform for natural language understanding. arXiv:1804.07461. Retrieved from https:\/\/arxiv.org\/abs\/1804.07461","DOI":"10.18653\/v1\/W18-5446"},{"key":"e_1_3_1_79_2","unstructured":"Binjie Wang Steffi Chern Ethan Chern and Pengfei Liu. 2024. Halu-J: Critique-based hallucination judge. arXiv: 240712943. Retrieved from https:\/\/arxiv.org\/abs\/2407.12943"},{"key":"e_1_3_1_80_2","unstructured":"Ben Wang and Aran Komatsuzaki. 2021. GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model. Retrieved from https:\/\/github.com\/kingoflolz\/mesh-transformer-jax"},{"key":"e_1_3_1_81_2","unstructured":"Peiyi Wang Lei Li Liang Chen Dawei Zhu Binghuai Lin Yunbo Cao Qi Liu Tianyu Liu and Zhifang Sui. 2023. Large language models are not fair evaluators. arXiv:2305.17926. Retrieved from https:\/\/arxiv.org\/abs\/2305.17926"},{"key":"e_1_3_1_82_2","unstructured":"Tianlu Wang Ping Yu Xiaoqing Ellen Tan Sean O\u2019Brien Ramakanth Pasunuru Jane Dwivedi-Yu Olga Golovneva Luke Zettlemoyer Maryam Fazel-Zarandi and Asli Celikyilmaz. 2023. Shepherd: A critic for language model generation. arXiv:2308.04592. Retrieved from https:\/\/arxiv.org\/abs\/2308.04592"},{"key":"e_1_3_1_83_2","unstructured":"Yizhong Wang Yeganeh Kordi Swaroop Mishra Alisa Liu Noah A. Smith Daniel Khashabi and Hannaneh Hajishirzi. 2022. Self-instruct: Aligning language model with self generated instructions. arXiv:2212.10560. Retrieved from https:\/\/arxiv.org\/abs\/2212.10560"},{"key":"e_1_3_1_84_2","first-page":"25","volume-title":"Proceedings of the Conference on Learning Theory","author":"Wang Yining","year":"2013","unstructured":"Yining Wang, Liwei Wang, Yuanzhi Li, Di He, and Tie-Yan Liu. 2013. A theoretical analysis of NDCG type ranking measures. In Proceedings of the Conference on Learning Theory. PMLR, 25\u201354."},{"key":"e_1_3_1_85_2","unstructured":"Yidong Wang Zhuohao Yu Zhengran Zeng Linyi Yang Cunxiang Wang Hao Chen Chaoya Jiang Rui Xie Jindong Wang Xing Xie et al. 2023. PandaLM: An automatic evaluation benchmark for LLM instruction tuning optimization. arXiv:2306.05087. Retrieved from https:\/\/arxiv.org\/abs\/2306.05087"},{"key":"e_1_3_1_86_2","unstructured":"Jason Wei. 2022. 137 Emergent Abilities of Large Language Models. Retrieved from https:\/\/www.jasonwei.net\/blog\/emergence"},{"key":"e_1_3_1_87_2","unstructured":"Yuxiang Wei Zhe Wang Jiawei Liu Yifeng Ding and Lingming Zhang. 2023. Magicoder: Source code is all you need. arXiv:2312.02120. Retrieved from https:\/\/arxiv.org\/abs\/2312.02120"},{"key":"e_1_3_1_88_2","unstructured":"Shijie Wu Ozan Irsoy Steven Lu Vadim Dabravolski Mark Dredze Sebastian Gehrmann Prabhanjan Kambadur David Rosenberg and Gideon Mann. 2023. BloombergGPT: A large language model for finance. arXiv:2303.17564. Retrieved from https:\/\/arxiv.org\/abs\/2303.17564"},{"key":"e_1_3_1_89_2","unstructured":"Can Xu Qingfeng Sun Kai Zheng Xiubo Geng Pu Zhao Jiazhan Feng Chongyang Tao and Daxin Jiang. 2023. Wizardlm: Empowering large language models to follow complex instructions. Retrieved from arXiv:2304.12244. Retrieved from https:\/\/arxiv.org\/abs\/2304.12244"},{"key":"e_1_3_1_90_2","doi-asserted-by":"crossref","unstructured":"Hongyang Yang Xiao-Yang Liu and Christina Dan Wang. 2023. FinGPT: Open-source financial large language models. arXiv:2306.06031. Retrieved from https:\/\/arxiv.org\/abs\/2306.06031","DOI":"10.2139\/ssrn.4489826"},{"key":"e_1_3_1_91_2","doi-asserted-by":"crossref","unstructured":"Rowan Zellers Ari Holtzman Yonatan Bisk Ali Farhadi and Yejin Choi. 2019. HellaSwag: Can a machine really finish your sentence? arXiv:1905.07830. Retrieved from https:\/\/arxiv.org\/abs\/1905.07830","DOI":"10.18653\/v1\/P19-1472"},{"key":"e_1_3_1_92_2","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing","author":"Zhang Qingru","year":"2023","unstructured":"Qingru Zhang, Dhananjay Ram, Cole Hawkins, Sheng Zha, and Tuo Zhao. 2023. Efficient long-range transformers: You need to attend more, but not necessarily at every layer. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing."},{"key":"e_1_3_1_93_2","unstructured":"Susan Zhang Stephen Roller Naman Goyal Mikel Artetxe Moya Chen Shuohui Chen Christopher Dewan Mona T. Diab Xian Li Xi Victoria Lin et al. 2022. OPT: Open pre-trained transformer language models. arXiv:2205.01068. Retrieved from https:\/\/arxiv.org\/abs\/2205.01068"},{"key":"e_1_3_1_94_2","unstructured":"Xinghua Zhang Bowen Yu Haiyang Yu Yangyu Lv Tingwen Liu Fei Huang Hongbo Xu and Yongbin Li. 2023. Wider and deeper LLM networks are fairer LLM evaluators. arXiv:2308.01862. Retrieved from https:\/\/arxiv.org\/abs\/2308.01862"},{"key":"e_1_3_1_95_2","unstructured":"Zhongwang Zhang and Zhi-Qin John Xu. 2023. Loss spike in training neural networks. arXiv:2305.12133. Retrieved from https:\/\/arxiv.org\/abs\/2305.12133"},{"key":"e_1_3_1_96_2","doi-asserted-by":"publisher","DOI":"10.1145\/3240323.3240374"},{"key":"e_1_3_1_97_2","first-page":"46595","article-title":"Judging LLM-as-a-judge with MT-bench and chatbot arena","volume":"36","author":"Zheng Lianmin","year":"2024","unstructured":"Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. 2024. Judging LLM-as-a-judge with MT-bench and chatbot arena. In Advances in Neural Information Processing Systems, Vol. 36 (2024), 46595\u201346623.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_98_2","unstructured":"Lianmin Zheng Wei Lin Chiang Ying Sheng Siyuan Zhuang Zhanghao Wu Yonghao Zhuang Lin Zi Zhuohan Li Dacheng Li Eric P. Xing et al. 2023. Judging LLM-as-a-judge with MT-bench and chatbot arena. arXiv:2306.05685. Retrieved from https:\/\/arxiv.org\/abs\/2306.05685"},{"key":"e_1_3_1_99_2","unstructured":"Deyao Zhu Jun Chen Xiaoqian Shen Xiang Li and Mohamed Elhoseiny. 2023. MiniGPT-4: Enhancing vision-language understanding with advanced large language models. arXiv:2304.10592. Retrieved from https:\/\/arxiv.org\/abs\/2304.10592"},{"key":"e_1_3_1_100_2","unstructured":"Lianghui Zhu Xinggang Wang and Xinlong Wang. 2023. Judgelm: Fine-tuned large language models are scalable judges. arXiv:2310.17631. Retrieved from https:\/\/arxiv.org\/abs\/2310.17631"}],"container-title":["ACM Transactions on Intelligent Systems and Technology"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3763000","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,28]],"date-time":"2026-04-28T14:41:43Z","timestamp":1777387303000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3763000"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,4,28]]},"references-count":99,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2026,8,31]]}},"alternative-id":["10.1145\/3763000"],"URL":"https:\/\/doi.org\/10.1145\/3763000","relation":{},"ISSN":["2157-6904","2157-6912"],"issn-type":[{"value":"2157-6904","type":"print"},{"value":"2157-6912","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,4,28]]},"assertion":[{"value":"2024-03-22","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-05-08","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-04-28","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}