A tutorial has been developed to evaluate the performance of multimodal vision models using PerceptionBench, a benchmark that assesses fine-grained visual perception capabilities across various tasks. The evaluation workflow involves configuring a suitable environment, installing necessary libraries, and loading a balanced dataset. The PerceptionBench benchmark measures tasks such as optical character recognition, object counting, and depth understanding. This evaluation process aims to provide a comprehensive assessment of multimodal vision models. The accuracy and reliability of these models in real-world applications depend on their performance in such benchmarks.