Research

AI and scientific progress

Network diagram linking a focal paper to its references, with nodes coloured by whether each reference is real or generated by GPT-4.
Citation graphs of a paper’s real and GPT-4-generated references. NAACL Findings 2025.

Large language models now read papers, suggest references, review manuscripts, and propose research methods. We ask how far they can speed up science and innovation, and what that does to the direction science takes. To find out, we check how well they know the scientific literature, how their references and methods compare with those written by scientists, and how peer review behaves when authors push back. This is the science of science, or metascience, with language models as both tool and subject.

How language models reason

Diagram of four access levels: the final response; plus a self-summary; plus the reasoning trace; plus provider-side features.
Four levels of access to a model’s run, each showing more of its work. Preprint, 2026.

Large language models often work through a problem step by step before they answer. We study what those steps tell us: whether longer thinking leads to better answers, which words in the steps hint that an answer is right or wrong, and how much of that thinking a user actually gets to see, which matters for transparency and AI safety.

Testing what AI can do

Bar chart: GPT-4o, o1, Claude 3.5 Sonnet, and Gemini 1.5 score between about 50 and 95 percent on GPQA, MATH, and MMLU, but under 10 percent on Humanity's Last Exam.
Models score high on older benchmarks and low on Humanity’s Last Exam. Figure: HLE consortium.

Claims about what AI can do rest on tests, and tests can mislead. We study when a test stops measuring what it claims to, and how to tell hard questions from easy ones without knowing the answers. Reliable evaluation is one foundation of AI safety.

Earlier work

Neural network diagram: gradient-based inverse design updates only the input layer, producing a canonical input for a preferred output, shown as histograms of input features.
LUCID designs the inputs a model prefers, to expose unfair treatment. AAAI 2023.

Before language models, I worked on two problems. The first was making algorithmic decisions transparent enough to find where they treat people unfairly. The second was turning text, such as daily news, into numbers that help predict economic phenomena, from consumer confidence to policy uncertainty in Belgium.

All papers

The full publication record is on Google Scholar and ORCID.

Team

PhD students

Postdoctoral researchers

We work closely with Vincent Ginis and the wider Data Analytics Lab.
Prospective PhD students and postdocs: openings are announced on this page.