1. Define the change
A window can appear in the right place while the model remains wrong. Its class, host relationship or opening may fail to match the intended edit. That is the practical question raised by BIM-Edit, a research benchmark for language-directed changes to IFC building models.
2. Read the benchmark
The preprint was first submitted on 18 June 2026 and revised on 23 June. It contains 324 editing tasks and evaluates geometry, semantics and topology separately. Among the tested systems, the highest mean score was about 49.5%, and no model fully solved more than 3.4% of tasks under the paper's strict criterion. Those are benchmark-specific results. BIM-Edit research record
3. Understand the test scope
The authors also describe important limits: a minimal code-execution setup, a finite interaction budget and one human-authored reference result per task. The study excludes several areas, including MEP and detailed structural systems. Its scores are a baseline for that experiment rather than an upper bound for every production assistant. Full paper and limitations
4. Turn the three axes into a review brief
GAD proposes a small acceptance sheet for an office pilot. Before the AI acts, write a precise change request and retain an untouched model copy. Identify the target elements, the intended change and the parts that must remain fixed. Choose a task the team can independently verify. Keep the test away from live coordination until its consequences are understood.
5. Check geometry
For geometry, compare location, dimensions and affected neighbouring elements. Use model views and measurements rather than relying on the assistant's completion message. A good review includes a view that would expose the most likely failure, such as a section through an edited opening.
6. Check meaning
For meaning, inspect the object class and the project information required for the next task. If a wall looks correct but cannot participate in the expected schedule or classification workflow, record that as an incomplete result. State which properties are essential so that optional metadata does not obscure a consequential omission.
7. Check relationships
For relationships, examine the dependencies that make the edit useful: containment, hosting, adjacency and connections relevant to the task. Move or resize a representative object in the normal authoring workflow and check the result. This additional revision helps reveal whether the output is maintainable.
8. Include a preservation check
Compare the surrounding model before and after. Look for objects or properties changed outside the requested scope. Have a reviewer who did not generate the result inspect the difference where practical. A concise change log should distinguish intended edits, necessary dependent edits and accidental alterations.
9. Record the decision
Accept, repair or reject the result using criteria chosen before the run. Track review effort, failed attempts and recovery alongside any speed benefit. The useful lesson is a disciplined evaluation method: a successful pilot leaves an intelligible building model and an auditable decision, with responsibility for design and technical approval still assigned to qualified people.
Worked review scenario
Use a small known IFC model and one independently checkable edit; inspect geometry, object meaning, relationships and unrequested changes.
When the pilot is ready to accept
The review criteria and unresolved limits remain explicit.
Another team member can trace the inputs and reproduce the acceptance decision.
The requested change is present, relevant IFC object relationships remain consistent, and all changes outside the agreed scope are resolved or rejected.