Multimodal Benchmark for Architecture & Civil Engineering

MMArch

Benchmarking Multimodal Reasoning Grounded in Architectural Evidence

Anonymous Authors

Overview

Overview of MMArch, including representative tasks, error distribution, and the human-model performance gap.

Figure 1. MMArch evaluates whether models can perceive distributed visual evidence, identify the governing engineering principle, and apply it to reach a conclusion.

Abstract

Multimodal large language models (MLLMs) perform strongly on engineering imagery, yet existing benchmarks mostly test drawing recognition, information extraction, or compliance checking, leaving open whether models can combine distributed visual evidence with engineering principles to reach a conclusion. We introduce MMArch, a benchmark for architecture and civil engineering spanning ten subdomains and built entirely from figures in peer-reviewed papers. Its 1,212 short-answer items are produced by a decoupled planner-writer pipeline and validated through automated screening, a blind adversarial audit, and expert review, so that answering requires perceiving the relevant evidence, identifying the governing principle, and applying it - not exploiting textual or single-figure shortcuts. Evaluating 18 open-weight and proprietary MLLMs against a domain-expert panel, we find a wide gap: the strongest open-source model attains about 30% and the best proprietary system 52%, while human experts reach 95%, more than forty points ahead. Our error analysis shows that failures concentrate in applying principles and combining evidence across figures rather than in locating it, pointing to substantial headroom for future research.

Benchmark Construction

The MMArch benchmark construction pipeline from peer-reviewed figures through question construction and quality control.

Figure 2. MMArch is constructed from authentic figures in peer-reviewed AEC papers using decoupled question-answer generation, shortcut screening, blind auditing, and expert review.