Publications
Grouped by year. For a complete and up-to-date list, see my Google Scholar profile. For representative work organised by research direction, see Research page.
Preprints
- GlobeAudio: A Multilingual Multicultural Benchmark for Naturalistic Evaluation of Large Audio-Language Models
- Beyond Alignment: Value Diversity as a Collective Property in Multicultural Agent Systems
- Agentic World Modeling: Foundations, Capabilities, Laws, and Beyond
- SeaLLMs-Audio: Large Audio-Language Models for Southeast Asia
- Babel: Open Multilingual Large Language Models Serving Over 90% of Global Speakers
2026
- MoltNet: Understanding Social Behavior of AI Agents in the Agent-Native MoltBook
- MASim: Multilingual Agent-Based Simulation for Social Science
- Language of Thought Shapes Output Diversity in Large Language Models
- DR-Arena: an Automated Evaluation Framework for Deep Research Agents SAC Highlight Award at ACL 2026
- Pardon? Evaluating Conversational Repair in Large Audio-Language Models
- Difficulty–Diversity Collaborative Filtering for Data-Efficient LLM Fine-Tuning
- PEAR: Phase Entropy Aware Reward for Efficient Reasoning
- SpaCE-Eval: A Benchmark for Real-World Multi-Modal Reasoning
- RAVENEA: A Benchmark for Multimodal Retrieval-Augmented Visual Culture Understanding
2025
- The Emergence of Abstract Thought in Large Language Models Beyond Any Language
- The Rise of Parameter Specialization for Knowledge Storage in Large Language Models
- Reframe Your Life Story: Interactive Narrative Therapist and Innovative Moment Assessment with Large Language Models
- Assessing Judging Bias in Large Reasoning Models: An Empirical Study
- Disentangling Language and Culture for Evaluating Multilingual Large Language Models
- Auto-Arena: Automating LLM Evaluations with Agent Peer Battles and Committee Discussions
- Pruning General Large Language Models into Customized Expert Models
- JsonTuning: Towards Generalizable, Robust, and Controllable Instruction Tuning
- Understanding and Enhancing Safety Mechanisms of LLMs via Safety-Specific Neuron
- SeaLLMs 3: Open Foundation and Chat Multilingual Large Language Models for Southeast Asian Languages
- Is Translation All You Need? A Study on Solving Multilingual Tasks with Large Language Models
- AdaMergeX: Cross-Lingual Transfer with Large Language Models via Adaptive Adapter Merging
- Zero-to-Strong Generalization: Eliciting Strong Capabilities of Large Language Models Iteratively without Gold Labels
2024
- How do Large Language Models Handle Multilingualism?
- On the Multi-turn Instruction Following for Conversational Web Agents
- SeaLLMs — Large Language Models for Southeast Asia
- Multilingual Jailbreak Challenges in Large Language Models
- Plug-and-Play Policy Planner for Large Language Model Powered Dialogue Agents
- Sentiment Analysis in the Era of Large Language Models: A Reality Check
2023
- M3Exam: A Multilingual, Multimodal, Multilevel Benchmark for Examining Large Language Models
- From Clozing to Comprehending: Retrofitting Pre-trained Language Model to Pre-trained Machine Reader
- SOUL: Towards Sentiment and Opinion Understanding of Language
- Knowledge-enhanced Mixed-initiative Dialogue System for Emotional Support Conversations
- Product Question Answering in E-Commerce: A Survey
- Bidirectional Generative Framework for Cross-domain Aspect-based Sentiment Analysis
- Easy-to-Hard Learning for Information Extraction
- Zero-Shot Text Classification via Self-Supervised Tuning
- A Unified Multi-task Learning Framework for Multi-goal Conversational Recommender Systems
2022 and earlier
- A Survey on Aspect-Based Sentiment Analysis: Tasks, Methods, and Challenges
- PACIFIC: Towards Proactive Conversational Question Answering over Tabular and Textual Data in Finance
- UniGDD: A Unified Generative Framework for Goal-Oriented Document-Grounded Dialogue
- User Satisfaction Estimation with Sequential Dialogue Act Modeling in Goal-oriented Conversational Systems
- Aspect Sentiment Quad Prediction as Paraphrase Generation
- Cross-lingual Aspect-based Sentiment Analysis with Aspect Term Code-Switching
- Aspect-based Sentiment Analysis in Question Answering Forums
- Towards Generative Aspect-Based Sentiment Analysis
- AnswerFact: Fact Checking in Product Question Answering
- Multi-hop Inference for Question-driven Summarization
- Answer Ranking for Product-Related Questions via Multiple Semantic Relations Modeling
- Review-guided Helpful Answer Identification in E-commerce
- Exploiting BERT for End-to-End Aspect-Based Sentiment Analysis