Wenqi Pei
HKU ECE MPhil

I am a MPhil student at The University of Hong Kong (HKU), supervised by Prof. Hongyang Du at NICE Lab. Prior to this, I completed my Bachelor's degree at the National University of Singapore (NUS) under the supervision of Prof. Bingsheng He. My research focuses on various aspects of Agentic AI, with a particular emphasis on Data Agent. I have also begun exploring multimodal Video-LLMs. I am open to collaborating with people who share similar research interests.

I am currently seeking PhD opportunities for Fall 2027. Please feel free to reach out if you see a potential research fit.


Education
  • The University of Hong Kong
    The University of Hong Kong
    Master of Philosophy
    Sep. 2025 - Present
  • National University of Singapore
    National University of Singapore
    Bachelor of Computing (Honours), Minor: Economics
    Aug. 2021 - Jul. 2025
Honors & Awards
  • Dean's List
    Sem2 2024-2025
  • Top Student for CS Research Methodology
    Sem1 2024-2025
News
2026
TempCloze accepted to EMNLP Finding 🎉
Aug 23
Released our survey on efficient tool-using LLM agents 📚
Aug 06
ROSE accepted to ACL Main 🎉
Apr 04
2025
NL2SQLBench accepted to VLDB 🎉
Dec 16
Feather-SQL accepted to IJCNLP-AACL Finding 🎉
Oct 25
Started graduate study at HKU 🎓
Sep 01
InterFeedback accepted to EMNLP Finding 🎉
Aug 20
Self-Monitoring Large Language Models accepted to ACM TOIS 🎉
Aug 17
Visited HKUST(GZ) DIAL Lab directed by Prof. Yuyu Luo for one month 🔬
Aug 01
Graduated from NUS with BComp (Honours) 🎓
Jul 15
Selected Publications (view all )
TempCloze: Can Video-LLMs Identify the Missing Middle?
TempCloze: Can Video-LLMs Identify the Missing Middle?

Wenqi Pei*, Hengyuan Zhao*, Yilai Liu*, Jiahao Meng, Han Chen, Ziyu Wang, Hongyang Du (* equal contribution)

EMNLP Finding 2026

We introduce TempCloze, a video cloze benchmark for evaluating visual temporal reasoning in Video-LLMs by asking them to identify the true missing middle between beginning and ending clips. Across 1,521 videos and 31 Video-LLMs, TempCloze reveals temporal alignment as the primary bottleneck.

TempCloze: Can Video-LLMs Identify the Missing Middle?

Wenqi Pei*, Hengyuan Zhao*, Yilai Liu*, Jiahao Meng, Han Chen, Ziyu Wang, Hongyang Du (* equal contribution)

EMNLP Finding 2026

We introduce TempCloze, a video cloze benchmark for evaluating visual temporal reasoning in Video-LLMs by asking them to identify the true missing middle between beginning and ending clips. Across 1,521 videos and 31 Video-LLMs, TempCloze reveals temporal alignment as the primary bottleneck.

How to Make Tool-Using LLM Agents Efficient? A Survey
How to Make Tool-Using LLM Agents Efficient? A Survey

Yunuo Hu*, Ziyang Cheng*, Wenqi Pei*‡, Miaomiao Li, Boyan Li, Yizhang Zhu, Yupeng Xie, Han Chen, Shizheng Hou, Yuyu Luo, Hongyang Du (* equal contribution, ‡ project lead)

Preprint 2026

We present a systematic survey and unified taxonomy of efficient tool-using LLM agents, organizing techniques across tool context, calling, reasoning, and execution. We synthesize shared mechanisms, limitations, open challenges, and future directions for reducing resource use while preserving task performance.

How to Make Tool-Using LLM Agents Efficient? A Survey

Yunuo Hu*, Ziyang Cheng*, Wenqi Pei*‡, Miaomiao Li, Boyan Li, Yizhang Zhu, Yupeng Xie, Han Chen, Shizheng Hou, Yuyu Luo, Hongyang Du (* equal contribution, ‡ project lead)

Preprint 2026

We present a systematic survey and unified taxonomy of efficient tool-using LLM agents, organizing techniques across tool context, calling, reasoning, and execution. We synthesize shared mechanisms, limitations, open challenges, and future directions for reducing resource use while preserving task performance.

ROSE: An Intent-Centered Evaluation Metric for NL2SQL
ROSE: An Intent-Centered Evaluation Metric for NL2SQL

Wenqi Pei*, Shizheng Hou*, Boyan Li*, Han Chen, Zhichao Shi, Yuyu Luo (* equal contribution)

ACL 2026

We introduce ROSE, an intent-centered evaluation metric for NL2SQL that uses an adversarial Prover-Refuter cascade to judge whether predicted SQL captures a user's intent. ROSE aligns more closely with expert assessments and supports a large-scale re-evaluation of 19 NL2SQL methods.

ROSE: An Intent-Centered Evaluation Metric for NL2SQL

Wenqi Pei*, Shizheng Hou*, Boyan Li*, Han Chen, Zhichao Shi, Yuyu Luo (* equal contribution)

ACL 2026

We introduce ROSE, an intent-centered evaluation metric for NL2SQL that uses an adversarial Prover-Refuter cascade to judge whether predicted SQL captures a user's intent. ROSE aligns more closely with expert assessments and supports a large-scale re-evaluation of 19 NL2SQL methods.

NL2SQLBench: A Modular Benchmarking Framework for LLM-Enabled NL2SQL Solutions [Experiment, Analysis & Benchmark]
NL2SQLBench: A Modular Benchmarking Framework for LLM-Enabled NL2SQL Solutions [Experiment, Analysis & Benchmark]

Shizheng Hou, Wenqi Pei, Nuo Chen, Quang-Trung Ta, Peng Lu, Beng Chin Ooi

VLDB 2026

We present NL2SQLBench, the first modular benchmarking framework for LLM-enabled NL2SQL approaches. We dissect NL2SQL systems into three modules: Schema Selection, Candidate Generation, and Query Revision. We review existing strategies and propose fine-grained metrics that systematically quantify module-level effectiveness and efficiency.

NL2SQLBench: A Modular Benchmarking Framework for LLM-Enabled NL2SQL Solutions [Experiment, Analysis & Benchmark]

Shizheng Hou, Wenqi Pei, Nuo Chen, Quang-Trung Ta, Peng Lu, Beng Chin Ooi

VLDB 2026

We present NL2SQLBench, the first modular benchmarking framework for LLM-enabled NL2SQL approaches. We dissect NL2SQL systems into three modules: Schema Selection, Candidate Generation, and Query Revision. We review existing strategies and propose fine-grained metrics that systematically quantify module-level effectiveness and efficiency.

InterFeedback: Unveiling Interactive Intelligence of Large Multimodal Models with Human Feedback
InterFeedback: Unveiling Interactive Intelligence of Large Multimodal Models with Human Feedback

Hengyuan Zhao*, Wenqi Pei*, Yifei Tao*, Haiyang Mei, Mike Zheng Shou (* equal contribution)

EMNLP Finding & ICLR Bi-Align (Oral) 2025

This work introduces InterFeedback, a novel approach to understanding and enhancing the interactive intelligence of large multimodal models through human feedback mechanisms. Our framework provides new insights into model behavior and interaction patterns.

InterFeedback: Unveiling Interactive Intelligence of Large Multimodal Models with Human Feedback

Hengyuan Zhao*, Wenqi Pei*, Yifei Tao*, Haiyang Mei, Mike Zheng Shou (* equal contribution)

EMNLP Finding & ICLR Bi-Align (Oral) 2025

This work introduces InterFeedback, a novel approach to understanding and enhancing the interactive intelligence of large multimodal models through human feedback mechanisms. Our framework provides new insights into model behavior and interaction patterns.

Feather-SQL: A Lightweight NL2SQL Framework with Dual-Model Collaboration Paradigm for Small Language Models
Feather-SQL: A Lightweight NL2SQL Framework with Dual-Model Collaboration Paradigm for Small Language Models

Wenqi Pei, Hailing Xu, Hengyuan Zhao, Shizheng Hou, Han Chen, Zining Zhang, Pingyi Luo, Bingsheng He

IJCNLP-AACL Finding & ICLR DL4C 2025

We present Feather-SQL, a novel framework for Natural Language to Structured Query Language (NL2SQL) that leverages dual-model collaboration paradigm specifically designed for small language models. Our approach achieves state-of-the-art (SOTA) results on the BIRD benchmark.

Feather-SQL: A Lightweight NL2SQL Framework with Dual-Model Collaboration Paradigm for Small Language Models

Wenqi Pei, Hailing Xu, Hengyuan Zhao, Shizheng Hou, Han Chen, Zining Zhang, Pingyi Luo, Bingsheng He

IJCNLP-AACL Finding & ICLR DL4C 2025

We present Feather-SQL, a novel framework for Natural Language to Structured Query Language (NL2SQL) that leverages dual-model collaboration paradigm specifically designed for small language models. Our approach achieves state-of-the-art (SOTA) results on the BIRD benchmark.

All publications