<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Preprint | Andrej Leban</title><link>https://andleb.netlify.app/publication-type/article/</link><atom:link href="https://andleb.netlify.app/publication-type/article/index.xml" rel="self" type="application/rss+xml"/><description>Preprint</description><generator>Hugo Blox Builder (https://hugoblox.com)</generator><language>en-us</language><lastBuildDate>Thu, 09 Jul 2026 00:00:00 +0000</lastBuildDate><image><url>https://andleb.netlify.app/media/icon_hu0b7a4cb9992c9ac0e91bd28ffd38dd00_9727_512x512_fill_lanczos_center_3.png</url><title>Preprint</title><link>https://andleb.netlify.app/publication-type/article/</link></image><item><title>CausalDS: Benchmarking Causal Reasoning in Data-Science Agents</title><link>https://andleb.netlify.app/publication/causalds/</link><pubDate>Thu, 09 Jul 2026 00:00:00 +0000</pubDate><guid>https://andleb.netlify.app/publication/causalds/</guid><description>&lt;p>The paper introduces a benchmark &lt;em>generator&lt;/em> for causal reasoning in agentic data-science workflows: every benchmark instance is synthetically generated rather than curated. In it, we:&lt;/p>
&lt;ul>
&lt;li>Generate fully synthetic &lt;em>scenes&lt;/em> pairing a sampled causal graph and structural causal model with tabular data and natural-language story;&lt;/li>
&lt;/ul>
&lt;ul>
&lt;li>Add an optional &lt;em>observation layer&lt;/em> that varies the data-analysis difficulty without changing the causal side;&lt;/li>
&lt;/ul>
&lt;ul>
&lt;li>Derive tasks spanning all three rungs of Pearl&amp;rsquo;s hierarchy which include non-answerable questions (i.e., testing abstention);&lt;/li>
&lt;/ul>
&lt;ul>
&lt;li>Make the benchmark fully parameterizable, so that it can be tailored to specific evaluation goals or grounded in empirical distributions obtained from real-world corpora.&lt;/li>
&lt;/ul>
&lt;p>We thus test the agents along five separate axes of &lt;em>agentic causal data science&lt;/em>: symbolic causal reasoning, data science, uncertainty quantification, abstention, and tool use/coding.&lt;/p>
&lt;p>We evaluate six contemporary agents — Claude Opus 4.8, Gemini 3.1 Pro, and GPT-5.5, along with the open-weight Qwen 3.6 35B, Kimi K2.6, and Gemma 4 26B — on a 100-task exam with realistically grounded composition. Claude Opus 4.8 leads overall, and the evaluated capabilities dissociate: symbolic causal reasoning is largely mastered across the field, while abstention, uncertainty quantification, and coding/tool-use/reasoning efficiency separate the models.&lt;/p>
&lt;p>The &lt;a href="https://huggingface.co/papers/2607.08093" target="_blank" rel="noopener">paper&lt;/a> was featured in &lt;a href="https://huggingface.co/papers/date/2026-07-10" target="_blank" rel="noopener">Hugging Face Daily Papers on July 10, 2026&lt;/a>.&lt;/p></description></item><item><title>A Bayesian approach to translators' reliability assessment</title><link>https://andleb.netlify.app/publication/bayesiantranslators/</link><pubDate>Tue, 12 Apr 2022 00:00:00 +0000</pubDate><guid>https://andleb.netlify.app/publication/bayesiantranslators/</guid><description>&lt;p>This grew out of the work I did as a Machine Learning intern at &lt;a href="https://translated.com" target="_blank" rel="noopener">Translated&lt;/a>
in Rome during the summer of 2021.&lt;/p>
&lt;p>Using sparse translation-quality annotations, we develop a Bayesian model that separates translation difficulty, translator skill, and reviewer behavior, including reviewer strictness and consistency.
We specifically use fat-tailed zero-inflated distributions to model the phenomenon where reviewers are likely to skim translations and assign zero errors; if they, however, &amp;ldquo;decide to look&amp;rdquo;, the error distribution becomes fat-tailed.&lt;/p>
&lt;p>Some of this work anticipates the more recent LLM eval literature, which has expanded dramatically since.&lt;/p></description></item></channel></rss>