Researcher, Reasoning Post-Training Current
Reasoning enhancement and post-training (SFT, RLHF, RLVR) for large language models: reward and evaluation design, reasoning-failure analysis, and large-scale experiments.
Reliable and controllable reasoning in language models and their agents
I work on the safety and reliability of large language models and the agents built on them — using uncertainty quantification to ask whether the training that makes them reason better also keeps them honest and calibrated, and to control them at inference time. MSc researcher at MBZUAI (advised by Timothy Baldwin and Artem Shelmanov); reasoning post-training in the Yandex Alice AI team; recently a research visit at the UKP Lab, TU Darmstadt with Iryna Gurevych. Background in applied mathematics at MIPT.
Open to PhD positions in technical AI safety · 2027/28 entry
My aim is large language models — and the agents built on them — that are safe and reliable. I approach this through the two levers I've worked on most: strengthening how LLMs reason, and applying uncertainty quantification to make their behavior more trustworthy, at both training time and inference time.
When should a model think longer, stop, verify, abstain, or escalate? I allocate inference-time compute by uncertainty rather than fixed heuristics, and study controllers that trade off accuracy, calibration, cost, and risk.
Does reasoning-oriented RL (RLHF, RLVR) buy real reliability, or reward-hack the verifier? I study how post-training reshapes calibration, and how to preserve honest uncertainty through aggressive fine-tuning.
Uncertainty compounds across long-horizon tool use. I work on propagating it across steps and tools, catching cascading errors before irreversible actions, and monitoring reasoning and coding agents.
Underpinned by uncertainty-quantification methods for LLMs (contributor to lm-polygraph) and UQ for diffusion language models.
Looking ahead. For a PhD I want to unify these threads into one question: how do we keep language models honest about their uncertainty as they are trained to reason and act — and use that uncertainty to control them at inference time? A concrete first step: study whether reasoning-oriented RL (RLHF / RLVR) preserves or erodes calibration, then build uncertainty-aware controllers that decide when a model should think longer, verify, abstain, or hand off — so that more capable reasoning also becomes more trustworthy and monitorable.
* equal contribution · full list on Google Scholar →
In preparation — ICLR 2027
Reasoning enhancement and post-training (SFT, RLHF, RLVR) for large language models: reward and evaluation design, reasoning-failure analysis, and large-scale experiments.
Uncertainty quantification and calibration for LLM agents (with Iryna Gurevych): scaling an SFT + RL pipeline with reward and evaluation infrastructure toward calibrated, trustworthy behavior.
Built TargetOS, a B2B LLM-agent platform: a cooperating multi-agent system (Creator, Corrector, Analyst, Superagent), LLM-as-a-judge evaluation, RAG optimization, and agent memory.
GPT-based shopping assistant, ranking models (CatBoost), and BERT search and product scoring distilled into production (NDCG 0.44 → 0.72).
Native desktop-client features and network-performance work.
MSc, Natural Language Processing. Fully funded UAE scholarship, 4.0/4.0. Advisors: Timothy Baldwin, Artem Shelmanov.
MSc, Finance (concurrent). Quantitative finance and statistical / mathematical methods for investment.
BSc, Applied Mathematics & Informatics. Thesis: Federated RLHF (BRAIn Lab, adv. A. Beznosikov), 4.6/5.0.
Graduate program in ML, NLP, and computer science.
Workshop on Uncertainty-Aware NLP (co-located with EMNLP 2026) — peer review of submissions on uncertainty estimation and calibration for NLP.
Based in Abu Dhabi, UAE. Open to PhD positions in technical AI safety (2027/28 entry) and to research collaborations on uncertainty, reasoning, and reliable LLM agents.