Researchers at Xiaomi have released a new benchmark for evaluating object removal models, which highlights the limitations of current metrics used to assess their performance. The benchmark reveals that popular metrics often fail to accurately rank the outputs of object removal models, which have improved significantly in recent years. This discrepancy is attributed to the inherently ill-posed nature of the object removal task, where multiple solutions can be considered correct. The findings suggest that the development of new, perception-aligned metrics is necessary to accurately evaluate object removal models. This matters because accurate evaluation of AI models is crucial for their adoption in real-world applications.