Research
AI and scientific progress
Large language models now read papers, suggest references, review manuscripts, and propose research methods. We ask how far they can speed up science and innovation, and what that does to the direction science takes. To find out, we check how well they know the scientific literature, how their references and methods compare with those written by scientists, and how peer review behaves when authors push back. This is the science of science, or metascience, with language models as both tool and subject.
- How deep do large language models internalize scientific literature and citation practices? Quantitative Science Studies 2026
- Dataset artefacts can partially drive the measured decline in disruption Nature 2026
- Structurally human, semantically biased: Detecting LLM-generated references with embeddings and GNNs ICLR 2026
- Large language models reflect human citation patterns with a heightened citation bias NAACL Findings 2025
- Thinking like a scientist? A structural study of LLM-generated research methods arXiv 2026
- Rebuttals move peer-review scores, but initial-review structure bounds the movement arXiv 2026
- Early evidence of vibe-proving with consumer LLMs: A case study on spectral region characterization with ChatGPT-5.2 (Thinking) arXiv 2026
How language models reason
Large language models often work through a problem step by step before they answer. We study what those steps tell us: whether longer thinking leads to better answers, which words in the steps hint that an answer is right or wrong, and how much of that thinking a user actually gets to see, which matters for transparency and AI safety.
- Lexical hints of accuracy in LLM reasoning chains Scientific Reports 2026
- The relationship between reasoning and performance in large language models: o3 (mini) thinks harder, not longer Scientific Reports 2026
- Probing the trajectories of reasoning traces in large language models NeurIPS 2026
- How much does a reasoning summary reveal? An observability ladder for large language models arXiv 2026
Testing what AI can do
Claims about what AI can do rest on tests, and tests can mislead. We study when a test stops measuring what it claims to, and how to tell hard questions from easy ones without knowing the answers. Reliable evaluation is one foundation of AI safety.
Before language models, I worked on two problems. The first was making algorithmic decisions transparent enough to find where they treat people unfairly. The second was turning text, such as daily news, into numbers that help predict economic phenomena, from consumer confidence to policy uncertainty in Belgium.
- LUCID: Exposing algorithmic bias through inverse design AAAI 2023
- Daily news sentiment and monthly surveys: A mixed-frequency dynamic factor model for nowcasting consumer confidence International Journal of Forecasting 2023
- Econometrics meets sentiment: An overview of methodology and applications Journal of Economic Surveys 2020
- The Economic Policy Uncertainty index for Flanders, Wallonia and Belgium BFW digitaal 2020
All papers
The full publication record is on Google Scholar and ORCID.
Team
PhD students
Postdoctoral researchers
We work closely with Vincent Ginis and the wider Data Analytics Lab.
Prospective PhD students and postdocs: openings are announced on this page.
