Portrait of Adrian Bulat

Adrian Bulat

Adrian Bulat

Principal Research ScientistSamsung AI Center Cambridge

Adrian Bulat is a Research Scientist at Samsung AI Cambridge. Previously, he received his PhD from the University of Nottingham where he worked with Dr. Georgios Tzimiropoulos as part of the Computer Vision Laboratory. His current research interests lie at the intersection of Computer Vision and Machine Learning.

Selected publications

Full publication list
  • 2026

    One Block, Multiple Depths: Recurrent Vision Transformers with Depth-Programmed Experts New

    Adrian Bulat, Yassine Ouali, Georgios Tzimiropoulos

    arXiv, 2026

    [abstract]

    reViT explores whether a single recurrent Transformer block can replace a full-depth vision encoder while preserving accuracy at comparable inference FLOPs. It shares attention across recurrent steps and constructs each step's feed-forward network by mixing a small shared bank of experts in weight space. A continuous normalized-depth coordinate controls the mixture, allowing different transformations at different depths. The paper evaluates supervised ImageNet-1k training and distillation from a DINOv2 teacher. A model with eight experts, trained using only the teacher's final output features, retains nearly all of the teacher's ImageNet linear-probe accuracy and transfers to classification, segmentation and depth prediction. Elastic-depth training enables one checkpoint to run at multiple tested depths. For deployment at a fixed depth, the recurrent block can also be expanded into a conventional dense graph, removing online routing and weight merging while increasing deployment storage.

    [cite]
    @article{bulat2026one,
      title={One Block, Multiple Depths: Recurrent Vision Transformers with Depth-Programmed Experts},
      author={Bulat, Adrian and Ouali, Yassine and Tzimiropoulos, Georgios},
      journal={arXiv preprint arXiv:2610.12448},
      year={2026},
      eprint={2610.12448},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2610.12448}
    }
  • 2026

    What Matters for Latent Reasoning with Flow Matching New

    Yassine Ouali, Adrian Bulat, Georgios Tzimiropoulos

    arXiv, 2026

    [abstract]

    FLaRe studies how language models can reason with continuous latent states and generate only their final answers. It combines a compact representation of symbolic reasoning, question-conditioned flow matching, and training on verified model-generated thoughts. The work evaluates whether these thoughts improve answers, support varied reasoning paths, admit faithful explanations, benefit from additional computation, and reduce inference cost. Experiments on arithmetic tasks examine the training choices needed to make latent reasoning effective.

    [cite]
    @article{ouali2026what,
      title={What Matters for Latent Reasoning with Flow Matching},
      author={Ouali, Yassine and Bulat, Adrian and Tzimiropoulos, Georgios},
      journal={arXiv preprint arXiv:2610.06666},
      year={2026},
      eprint={2610.06666},
      archivePrefix={arXiv},
      primaryClass={cs.LG},
      url={https://arxiv.org/abs/2610.06666}
    }
  • 2026

    UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models New

    Ioannis Maniadis Metaxas*, Adrian Bulat*, Alberto Baldrati*, Anestis Zaganidis, Yassine Ouali, Hyeonuk Kim, Georgios Tzimiropoulos

    European Conference on Computer Vision (ECCV), 2026

    [abstract]

    UltraViT addresses the cost of visual encoding when deploying large vision-language models on edge devices. Its pyramidal encoder combines different spatial mixing operations, with architectural choices guided by measured device latency rather than computational proxies alone. A two-stage pre-training procedure first transfers detailed spatial representations through dense distillation and then applies generative supervision from a frozen language model trained with mixed capacity. This procedure develops the semantic grounding needed for subsequent multimodal alignment and is reported to outperform contrastive and self-supervised alternatives. Experiments show that the combination of device-aware architecture design and generative training improves the efficiency and performance of vision encoding for large vision-language models, exceeding encoder-focused baselines while achieving approximately 1.7 times their on-device speed.

    [cite]
    @inproceedings{metaxas2026ultravit,
      title={{UltraViT}: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models},
      author={Metaxas, Ioannis Maniadis and Bulat, Adrian and Baldrati, Alberto and Zaganidis, Anestis and Ouali, Yassine and Kim, Hyeonuk and Tzimiropoulos, Georgios},
      booktitle={European Conference on Computer Vision (ECCV)},
      year={2026},
      eprint={2607.23373},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2607.23373}
    }
  • 2026

    Hierarchical Image Tokenization for Multi-Scale Image Super Resolution New

    Isma Hadji*, Enrique Sanchez*, Adrian Bulat, Brais Martinez, Georgios Tzimiropoulos

    International Conference on Machine Learning (ICML), 2026

    [abstract]

    This work introduces HIT, an approach to image super-resolution that adapts visual autoregressive modeling to produce several output resolutions in one forward pass. Conventional residual quantization does not ensure that intermediate token scales correspond to image scales, limiting earlier methods to a fixed output resolution. HIT instead builds a hierarchy of image tokens with overlap between scales, encouraging consistent representations at each resolution. Training also incorporates a direct preference optimization objective that favors the high-resolution target over its low-resolution input, using only the paired training images. Together, these changes support a 300-million-parameter model, compared with the billion-parameter VARSR baseline, without requiring external training data. The resulting system reports state-of-the-art super-resolution performance while reducing model size and providing outputs at multiple scales.

    [cite]
    @inproceedings{hadji2026hierarchical,
      title={Hierarchical Image Tokenization for Multi-Scale Image Super Resolution},
      author={Hadji, Isma and Sanchez, Enrique and Bulat, Adrian and Martinez, Brais and Tzimiropoulos, Georgios},
      booktitle={International Conference on Machine Learning (ICML)},
      year={2026},
      eprint={2605.14891},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2605.14891}
    }

Research

His current research interests lie at the intersection of Computer Vision and Machine Learning, with work conducted on topics such as efficient neural networks (via bit quantization, network binarization and compression) and human analysis (face alignment/recognition/super-resolution and human pose estimation).

Teaching

  • 2016–2018
    Teaching Assistant · University of Nottingham

    G52CPP: 2nd-year C++ Programming. G53SEC: 3rd-year Network Security.