SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models
Detects factual inconsistency by comparing multiple samples from a black-box generative model.
RESEARCH
Our research spans AI safety, specialized language models, reasoning, tool use, retrieval, speech, and evaluation. We turn frontier AI research into reliable systems built for real-world deployment.
Published at leading AI conferences






Detects factual inconsistency by comparing multiple samples from a black-box generative model.
Uses pairwise judgments from large language models to evaluate generated text without task-specific training data.
Generates multiple-choice questions to test whether a summary remains consistent with its source document.
Extends consistency checking into a comparative hallucination-ranking framework for multimodal foundation models.
Evaluates whether medical reasoning models remain correct across multiple-choice, open-ended, and ranked-list answer formats.
Introduces an open family of Thai-focused large language models and training resources.
Develops multilingual language models designed to broaden coverage across Southeast Asian languages.
Presents a rapid model-merging recipe for adding reasoning capability to a language-specific model.
Documents an open approach for adapting a Thai language model to structured reasoning tasks.
Provides a minimal open post-training recipe for building deployable sovereign language models.
Examines how program-of-thought reasoning transfers across languages and multilingual settings.
Grounds chain-of-thought generation in financial reasoning patterns drawn from domain expertise.
Improves role-play agents through automatic prompt optimization, concise dialogue behavior, and more reliable tool calls.
Aggregates multiple retrieval and reasoning paths to answer questions over large knowledge collections.
Adapts tool-agent-user evaluation to low-resource Southeast Asian languages and localized tool-use settings.
Tests generative large language models as post-processing systems for correcting automatic speech-recognition errors.
Introduces low-resource language adaptation and instruction-following methods for audio-language models.
Compares static benchmarks with interactive evaluation to expose gaps in large audio-model performance.
Presents a streaming Thai speech-recognition system built around a FastConformer-Transducer architecture.