DrawingVQA — benchmarking MLLMs on construction drawings
New dataset targets a blind spot: multimodal models struggle with construction drawings, which layer geometry, symbols, tables, annotations, and domain text in ways natural images don't.
DrawingVQA pairs 33 real “Issued for Construction“ drawings with 92 expert Q&A across three reasoning depths—perception, context, and domain expertise.
First benchmark to stress-test MLLMs on engineering workflows at scale.