Instance-Level Recognition and Generation Workshop at ECCV'26

Sep 8, 8:30am-12:45pm (CEST)

Workshop Location: Quality View Hotel - Stroget 1

Poster assignments

ILR+G 2026

The main focus of our workshop is on computer vision tasks that operate at instance-level, including both recognition (instance-level recognition - ILR) and generation (instance-level generation - ILG), denoted as ILR+G. More precisely, ILR+G is the task of identifying, comparing, and generating images of specific objects, scenes, or events.

This year, we will organize a call for papers, and host keynote talks by renowned speakers and invited paper talks from the main conference.

The 2026 Instance-Level Recognition and Generation (ILR+G) Workshop is a follow-up of seven successful editions of our previous workshops — the first two having focused only on landmark recognition (CVPRW18, CVPRW19), the following ones expanding to the domains of artworks and products (ECCVW20, ICCVW21), introducing the universal image embedding problem (ECCVW22, ECCVW24), and the latest one expanding the scope of our workshop to ILG and the potential synergy between ILG and ILR (ICCVW25).

Workshop Schedule

Welcome remarks

Sep. 8, 8:30am-8:40am (CEST)

Keynote 1

Richard Zhang
Steering Models with Notions of Similarity

Sep. 8, 8:40am-9:10am (CEST)

Keynote 2

Adam Harley
Getting to the Point: From Instances to Correspondence

Sep. 8, 9:35am-10:05am (CEST)

Poster session & Coffee break

Poster assignments
Sep. 8, 10:05am-11:15am (CEST)

Keynote 3

Benjamin Busam
Open-World Object Reasoning

Sep. 8, 12:10pm-12:40pm (CEST)

Closing & Award

Sep. 8, 12:40pm-12:45pm (CEST)

Keynote Speakers

Benjamin Busam

Technical University of Munich

Open-World Object Reasoning

Adam Harley

Meta Reality Labs

Getting to the Point: From Instances to Correspondence

Richard Zhang

Adobe Research

Steering Models with Notions of Similarity


Accepted Papers

Long Papers

  • READ THE OBJECT: A Dataset and Task for Text-Grounded Object Understanding in Real-World Images (Oral)
    Arpita Chowdhury, Hyojin Park, Hoang Le, Wei-Lun Chao, Munawar Hayat, Fatih Porikli, Shweta Mahajan
  • Instance-Level Image-Based Shape Retrieval via Pre-Aligned Multi-Modal Encoders and Hard Contrastive Learning (Oral)
    Paul Julius Kühn, Cedric Spengler, Saptarshi Neil Sinha, Michael Weinmann, Arjan Kuijper
  • Cross-Temporal Building-Instance Retrieval: A Benchmark for Disaster Response
    Michaela Areti Zervou, Grigorios Tsagkatakis, Panagiotis Tsakalides
  • Benchmarking Personalization Methods for Text-to-Image Generation: Dataset, Metrics, and a Comprehensive Evaluation Framework
    Kutay Özbay, Yalın Baştanlar
  • Short Papers

  • Neither Name nor Schema: Pairwise Instance Labels for Hairstyle
    Jordan Lim
  • Visual-Prompt Guided Wildlife Instance-Level Recognition
    Mufhumudzi, Jiahao Huo, Terence L. van Zyl, Fredrik Gustafsson
  • Foveate: A Training-Free, In-Context, Recursive Method for Instance Segmentation
    Robert Andreas Leist, Thiago S. Gouvêa, Daniel Sonntag
  • Invited Papers

  • FoundYou: A Unified Model for Personalized Segmentation and Retrieval (Oral)
    Gabriele Trivigno, Marcos Alfaro Perez, Claudia Cuttano, Gabriele Berton, Luis Payá, Carlo Masone
  • DreamEdit3D: Personalization of Multi-View Diffusion Models for 3D Editing (Oral)
    Jinxin Ai, Matthias Nießner, Ziya Erkoç
  • NearID: Identity Representation Learning via Near-identity Distractors (Oral)
    Aleksandar Cvejic, Rameen Abdal, Abdelrahman Eldesokey, Bernard Ghanem, Peter Wonka
  • GH-ESD: Grounded Hypothesis-Driven Error Slice Discovery for Instance-Level Vision Tasks
    Wei Zhang, Chaoqun Wang, Zixuan Guan, Ping Sheng Kao, Pengfei Zhao, Peng Wu, Sifeng He
  • AutoCompass: Accurate Visual Localization on Public Maps by Learning from Weak Labels
    Javier Tirado-Garín, Alan Paul, Shuai Chen, Axel Barroso-Laguna, Tommaso Cavallari, Daniyar Turmukhambetov, Victor Adrian Prisacariu, Eric Brachmann
  • Vulnerability of Privacy-Preserving Visual Localization against Diffusion-based Attacks
    Maxime Pietrantoni, Torsten Sattler, Gabriela Csurka
  • LoMa: Local Feature Matching Revisited
    David Nordström, Johan Edstedt, Georg Bökman, Jonathan Astermark, Anders Heyden, Viktor Larsson, Mårten Wadenbäck, Michael Felsberg, Fredrik Kahl
  • XYZ-IBD: Benchmarking Robust 6D Object Pose Estimation under Real-World Industrial Complexity
    Junwen Huang, Jiaqi Hu, Peter Yu, Slobodan Ilic, Martin Sundermeyer, Benjamin Busam
  • Pose Anything Anywhere: Model-free Object Poses from Arbitrary References
    Hongli Xu, Jiaqi Hu, Junwen Huang, Boyang Zhong, Peter KT Yu, Nassir Navab, Benjamin Busam, Slobodan Ilic
  • Beyond Pixel Mimicry: Disentangled Self-Similarity Rewards for Diverse Subject-Driven Generation
    Qian Wang, Zhenyu Li, Abdelrahman Eldesokey, Peter Wonka
  • Point2Pose: Occlusion-Recovering 6D Pose Tracking and 3D Reconstruction for Multiple Unknown Objects Via 2D Point Trackers
    Tzu-Yuan Lin, Ho Lee, Kevin Doherty, Yonghyeon Lee, Sangbae Kim
  • PS-MOT: Cultivating Instance Awareness from Point Seeds for Multi-Object Tracking
    Kai Luo, Fei Teng, Mengfei Duan, Wanjun Jia, Xu Wang, Hao Shi, Kunyu Peng, Zhiyong Li, Kailun Yang
  • Track4World: Feedforward World-centric Dense 3D Tracking of All Pixels
    Jiahao Lu, Jiayi Xu, Wenbo Hu, Ruijie Zhu, Chengfeng Zhao, Sai Kit Yeung, Ying Shan, Yuan Liu
  • Where and What: Long-Term Object Tracking in Egocentric Videos
    Jacob Chalk, Saptarshi Sinha, Dima Damen, Yannis Kalantidis, Diane Larlus
  • SOS!: A Streamlined Object-Conditional Transformer for Model-free Segmentation
    Jiaqi Hu, Junwen Huang, Hongli Xu, Peter KT Yu, Nassir Navab, Benjamin Busam, Slobodan Ilic
  • RoMa v2: Harder Better Faster Denser Feature Matching
    Johan Edstedt, David Nordström, Yushan Zhang, Georg Bökman, Jonathan Astermark, Viktor Larsson, Anders Heyden, Fredrik Kahl, Mårten Wadenbäck, Michael Felsberg
  • Online 3D Instance Segmentation at task-oriented granularity with Unposed Monocular Video
    Dong Wu, Baicheng Li, Yingdian Cao, Shunkai Zhou, Yiwen Lu, Hongbin Zha
  • MMLANDMARKS: a Cross-View Instance-Level Benchmark for Geo-Spatial Understanding
    Oskar Kristoffersen, Alba Reinders Sánchez, Morten Rieger Hannemose, Anders Bjorholm Dahl, Dim P Papadopoulos
  • Indexing Multimodal Language Models for Large-scale Image Retrieval
    Bahey Tharwat, Giorgos Kordopatis-Zilos, Pavel Suma, Ian Reid, Giorgos Tolias
  • ELViS: Efficient Visual Similarity from Local Descriptors that Generalizes Across Domains
    Pavel Suma, Giorgos Kordopatis-Zilos, Yannis Kalantidis, Giorgos Tolias
  • Call For Papers

    We call for novel and unpublished work in the format of long papers (14 pages excluding references) and short papers (4 pages excluding references). Papers should follow the ECCV proceedings style and will undergo double-blind peer review. Selected long papers will be invited for oral presentations; all accepted papers will be presented as posters. Only long papers will be published in the ECCV workshop proceedings. This year, financial awards will be given to the best papers, recognizing high-quality research with novel and interesting findings, and student support grants will be provided to eligible participants. All submissions will be handled electronically via the OpenReview submission system.

    Topics of interest include (but are not limited to)

  • instance-level object classification, detection, segmentation, and pose estimation
  • particular object and event retrieval
  • personalized image and video generation
  • cross-modal/multi-modal recognition at instance-level
  • other ILR tasks such as image matching, visual geo-localization, animal re-identification, copy detection, video tracking, moment retrieval
  • other ILR+G applications, datasets, and benchmarks

  • Even though tasks such as person and vehicle re-identification fall within the definition of ILR, we intentionally omit them from the list of topics, due to ethical and social implications. Submitted papers on those topics will be desk rejected.

    Important Dates

  • Submission deadline: June 26, 2026 July 3rd, 2026
  • Notification of acceptance: July 24, 2026
  • Camera-ready papers due: August 15, 2026

  • Student Support Grant

    To broaden participation, we are offering student support grants to help cover registration costs. Eligible students are encouraged to apply here.

  • Application deadline: July 31, 2026
  • Notification of decisions: August 4, 2026

  • Questions? Please reach out to us at ilr-workshop@googlegroups.com

    Organizers

    Andre Araujo

    Google DeepMind

    Bingyi Cao

    Google DeepMind

    Kaifeng Chen

    xAI

    Ondrej Chum

    Czech Technical University

    Noa Garcia

    The University of Osaka

    Guangxing Han

    Google DeepMind

    Giorgos Kordopatis-Zilos

    Czech Technical University (Primary Contact)

    Giorgos Tolias

    Czech Technical University

    Yankun Wu

    The University of Osaka

    Hao Yang

    Amazon

    Nikolaos-Antonios Ypsilantis

    Czech Technical University

    Xu Zhang

    Amazon

    The Microsoft CMT service was used for managing the peer-reviewing process for this conference. This service was provided for free by Microsoft and they bore all expenses, including costs for Azure cloud services as well as for software development and support.

    © 2026 ILR+G 2026

    We thank Jalpc for the jekyll template