Andrej Leban
Andrej Leban
Home
Research
Publications
Talks
Honors & Media
Teaching
Posts
Projects
Light
Dark
Automatic
Causal Inference
CausalDS: Benchmarking Causal Reasoning in Data-Science Agents
CausalDS is a
benchmark generator
for agentic causal data science: every problem is generated fresh in its entirety, with tasks spanning all three rungs of Pearl’s hierarchy that involve significant tool use. Exam composition is a free parameter: it can be tailored to a specific goal or grounded in real-world corpora. The benchmark jointly tests symbolic causal reasoning, data-science execution, uncertainty quantification, epistemic abstention, and coding/tool use.
Andrej Leban
,
Yuekai Sun
arXiv preprint
PDF
Cite
Code
arXiv
alphaXiv
Hugging Face dataset
New preprint: CausalDS
We have a new preprint out on benchmarking agentic causal data-science.
Last updated on Jul 9, 2026
1 min read
Approaching an unknown communication system ... accepted to Royal Society Open Science
Approaching an unknown communication system has been accepted to Royal Society Open Science.
Last updated on Jun 18, 2026
1 min read
A talk on our work on whale communication at the Simons Institute
A talk at the Simons Institute on using deep generative models to study sperm whale communication.
Last updated on Jul 11, 2023
1 min read
Approaching an unknown communication system by latent space exploration and causal inference
We propose a novel methodology -
Causal Disentanglement with Extreme Values (CDEV)
- to identify representations learned by GANs. When trained on raw whale communication, it finds - for the first time - specific acoustic attributes that might serve as carriers of meaning.
Gašper Beguš
,
Andrej Leban
,
Shane Gero
Royal Society Open Science
, 13: 250829
PDF
Cite
Code
Royal Society Open Science
New publication: Approaching an unknown communication system
We have a pre-print out based on our work with Project CETI, an initiative to decipher sperm whale communication using machine learning. In the paper, we combine approaches that help discover how known properties of human language are learned by generative models when trained on labeled speech audio data with methodology inspired by causal inference and apply it to the communication system of sperm whales, for which we do not have any such ground truth.
Last updated on Mar 20, 2023
1 min read
Cite
×