Killian
Steunou

Hi! I'm a PhD student in machine learning at Institut Polytechnique de Paris and Moments Lab, and these days I'm trying to find ways for models to understand the world through video. For me that means omni-modal models that see and hear at the same time, run efficiently, and keep working on footage that looks nothing like their training data. So far it has led to a way of picking the frames that matter [2]PEEK, BMVC 2026. Frame selection distilled from vision-language teachers. and a survey of what actually makes video LLMs cheaper [1]Why is video still so expensive? A survey of 125 inference-efficiency methods for video LLMs..

Before the PhD, I studied maths and statistics at Toulouse School of Economics and then did the MVA master at ENS Paris-Saclay, with research internships at Idemia, CLS and JoliBrain along the way. I care about open code and results that other people can reproduce.

News

older

Publications

  1. BMVC2026

    PEEK: Picking Essential Frames via Efficient Knowledge Distillation

    Killian Steunou, Anas Filali Razzouki, Mounîm A. El-Yacoubi, Khalil Guetari, Yannis Tevissen

Background

Work

Education

  • 2025–

    PhD in machine learning, Institut Polytechnique de Paris

  • 2024–25

    Master MVA, ENS Paris-Saclay

    Optimal transport, convex optimization, generative models.

  • 2023–24

    M1 applied maths and statistics, Toulouse School of Economics

  • 2022–23

    Exchange semester, University of Copenhagen

  • 2019–22

    Bachelor in maths and economics, Toulouse School of Economics

Projects

  • Video background removal

    Removes the background of a whole video with Mobile SAM.

    codedemo
  • joliGEN

    Contributed edge controls and SAM-based masking to JoliBrain's generative toolkit.

    code
  • Video object detection

    Zero-shot detection in videos with OWL-ViT and text prompts.

    code
  • Nail bite detection

    A macOS menu-bar app that notices nail biting through the webcam.

    codesite
  • Audio-visual transcription

    Subtitles for audio and video files with Whisper.

    codedemo
  • MathViz

    Interactive visualisations of maths, statistics and ML ideas.

    codedemo
coursework reports
  • Score-based generative networks for large-scale optimal transport

    SCONES reproduction.

    reportcode
  • Online test-time training with masked autoencoders

    TTT-MAE extension.

    reportcode
  • An end-to-end transformer for 3D object detection

    3DETR reproduction.

    report
  • Are generative classifiers more robust to adversarial attacks?

    Adversarial robustness.

    reportcode
  • Toxic gas characterisation under humidity shift

    Multi-task learning and adversarial domain adaptation.

    reportcode
  • Convergence of SGD with sliced Wasserstein losses

    Optimisation for generative modelling.

    reportcode

Writing

Contact

contact@killian-steunou.com

I'm always happy to talk about efficient video models, multimodal learning, evaluation, or a possible collaboration.