Andrej Leban
Andrej Leban
Home
Research
Publications
Talks
Honors & Media
Teaching
Posts
Projects
Light
Dark
Automatic
Preprint
CausalDS: Benchmarking Causal Reasoning in Data-Science Agents
CausalDS is a
benchmark generator
for agentic causal data science: every problem is generated fresh in its entirety, with tasks spanning all three rungs of Pearl’s hierarchy that involve significant tool use. Exam composition is a free parameter: it can be tailored to a specific goal or grounded in real-world corpora. The benchmark jointly tests symbolic causal reasoning, data-science execution, uncertainty quantification, epistemic abstention, and coding/tool use.
Andrej Leban
,
Yuekai Sun
arXiv preprint
PDF
Cite
Code
arXiv
alphaXiv
Hugging Face dataset
A Bayesian approach to translators' reliability assessment
Modeling the translation and review processes as zero-inflated fat-tailed distributions, we show how to extract useful information on translators’ reliability with as little as one review per translation.
Marco Miccheli
,
Andrej Leban
,
Andrea Tacchella
,
Andrea Zaccaria
,
Dario Mazzilli
,
Sébastien Bratières
arXiv preprint
PDF
Cite
arXiv
Cite
×