Daniel Marta

Postdoc, Learning & Adaptive Systems Group, ETH Zürich

I’m a postdoc in the Learning & Adaptive Systems Group at ETH Zürich, supervised by Prof. Andreas Krause, since September 2025. My research aims to enable autonomous systems, physical or virtual, to operate in the real world through large-scale multimodal alignment, uncertainty-aware adaptation, and human feedback.

Before joining ETH, I was a Ph.D. student at RPL (Robotics, Perception and Learning), KTH Royal Institute of Technology, and a research visitor at Yale University in the Interactive Machines Group. I am also a member of the WASP doctoral program.

Previously, I researched combinatorial optimization using machine learning at TU Delft in the Air Transport and Operations department, and interned at Latent Logic, a self-driving car startup later acquired by Waymo.

news

Jun 2026MOSAIC published at ICRA 2026, reframing preference-based RL as multi-objective learning from language explanations. Project page.
Jun 2026Uncertainty Quantification for Flow-Based Vision-Language-Action Models accepted to the RSS 2026 Diff4RL Workshop, introducing VFD for failure detection and SAVE for active VLA fine-tuning. Project page.
May 2026Awarded a compute grant from Swiss AI for CSCS Alps extreme-scale HPC access.
Apr 2026Reinforcement Learning via Self-Distillation accepted to ICML 2026. ICML page · website · code.
Mar 2026SDPO accepted to the ICLR 2026 Workshop on Scaling Post-training for LLMs (SPOT).
Sep 2025Joined the Learning & Adaptive Systems Group at ETH Zürich as a postdoc.
May 2025FLoRA published at ICRA 2025, adapting reward functions with low-rank style parameters for sample-efficient preference-based RL. Project page.
Nov 2024Awarded a research visit to Japan focused on reinforcement learning and industrial AI applications for manufacturing, including automotive manufacturing with Toyota Motor and Toyota Industries Corporation, manufacturing and computer vision with Toshiba, robotics with Hitachi, and AI/finance with DMG Mori, with the Swedish Embassy in Tokyo and Nagoya.
2024POLITE was a finalist for the ICRA 2024 Best Conference Paper Award, Best Student Paper Award, and Best Paper Award on Human-Robot Interaction.
Apr 2024Started a research exchange visit at Yale University, New Haven CT, USA, with the Interactive Machines Group led by Marynel Vázquez.
2024SEQUEL and POLITE published at ICRA 2024.
2024Human-Centric Autonomous Systems With LLMs for User Command Reasoning received the Best Student Paper Award at the WACV 2024 LLVM-AD Workshop.
2024PREDILECT published at HRI 2024, using zero-shot language-model reasoning for preference-based RL. ACM page.
2023VARIQuery presented at IROS 2023 in Detroit, Michigan.
2023Aligning Human Preferences with Baseline Objectives in Reinforcement Learning presented at ICRA 2023 in London, United Kingdom.
2021Human-feedback shield synthesis for perceived safety in deep reinforcement learning published in IEEE Robotics and Automation Letters.

latest paper

project page
RSS Diff4RL 2026

Uncertainty Quantification for Flow-Based Vision-Language-Action Models

We introduce Velocity-Field Disagreement (VFD) for epistemic uncertainty in flow-based VLAs and SAVE, an uncertainty-guided active fine-tuning framework. On LIBERO, SAVE reduces expert data needs by at least 22% while improving failure awareness and adaptation.

VFD and SAVE method overview
VFD + SAVE overview
Failure detection with VFD
failure detection
SAVE active fine-tuning results
active fine-tuning

selected publications

view all
  1. RSS Diff4RL 2026 Uncertainty Quantification for Flow-Based Vision-Language-Action Models thumbnail
    2026

    Uncertainty Quantification for Flow-Based Vision-Language-Action Models

    Ralf Römer, Maximilian Seeliger, Saida Liu, Ben Sturgis, Marco Bagatella, Daniel Marta, Andreas Krause, and Angela P. Schoellig

    Accepted to the RSS 2026 Diff4RL Workshop; arXiv:2606.18043, 2026

    Introduces Velocity-Field Disagreement (VFD) for epistemic uncertainty in flow-based VLAs and SAVE, an uncertainty-guided active fine-tuning framework that needs at least 22% fewer expert demonstrations than baselines.

  2. ICRA 2026 MOSAIC: Multi-objective Optimization from Zero-Shot Language Reasoning in Preference-based RL thumbnail
    2026

    MOSAIC: Multi-objective Optimization from Zero-Shot Language Reasoning in Preference-based RL

    Daniel Marta*, Simon Holk*, et al.

    In IEEE International Conference on Robotics and Automation (ICRA), 2026

    Reframes preference-based RL as multi-objective learning: language explanations are parsed into objective-specific labels, weights, and highlights for scalarized policy optimization.

  3. ICML + SPOT 2026 Reinforcement Learning via Self-Distillation thumbnail
    2026

    Reinforcement Learning via Self-Distillation

    Jonas Hübotter, Frederike Lübeck*, Lejs Behric*, Anton Baumann*, Marco Bagatella, Daniel Marta, Ido Hakimi, Idan Shenfeld, Thomas Kleine Buening, Carlos Guestrin, and Andreas Krause

    In International Conference on Machine Learning (ICML), 2026; also accepted to the ICLR 2026 Workshop on Scaling Post-training for LLMs (SPOT)

    Introduces Self-Distillation Policy Optimization (SDPO), converting rich environment feedback into dense self-distillation signals for more sample-efficient RL with language models.

  4. ICRA 2025 FLoRA: Sample-Efficient Preference-based RL via Low-Rank Style Adaptation of Reward Functions thumbnail
    2025

    FLoRA: Sample-Efficient Preference-based RL via Low-Rank Style Adaptation of Reward Functions

    Daniel Marta*, Simon Holk*, Miguel Vasco, Jens Lundell, Timon Homberger, Finn Busch, Olov Andersson, Danica Kragic, and Iolanda Leite

    In IEEE International Conference on Robotics and Automation (ICRA), 2025

    Adapts reward functions with low-rank style parameters to support preference-based robot adaptation from limited feedback while mitigating catastrophic reward forgetting.

  5. ICRA 2024 SEQUEL: Semi-Supervised Preference-based RL with Query Synthesis via Latent Interpolation thumbnail
    2024

    SEQUEL: Semi-Supervised Preference-based RL with Query Synthesis via Latent Interpolation

    Daniel Marta*, Simon Holk*, Christian Pek, Jana Tumova, et al.

    In IEEE International Conference on Robotics and Automation, 2024

    Improves preference-learning sample efficiency by augmenting human feedback with synthesized preference queries from latent interpolation.