GAD Editorial · Published · · English study · Chinese summary
Before asking an image model to select a window, decide what the word means in the task. Does the selection include the frame, the reveal, the reflected sky or only the visible glazing? A facade study becomes difficult to compare when each image answers that question differently.
Meta's SAM 2 offers promptable segmentation of images and video. A user can indicate an object with a point, box or mask, then refine the prediction with additional prompts. Its video workflow carries information across frames. The research model, code and demo are public, but these capabilities do not establish an architectural condition-assessment or measurement system. Official project
The paper was first submitted in August 2024. The following is an evergreen editorial application of that work, not a claim that Meta has launched a facade-survey product. GAD has not tested SAM 2 on a project image set. Paper and version history
Write a selection rule before creating masks
Choose a narrow visual question, such as preparing a consistent set of window regions for a comparative facade study. Define the visible boundary and exclusions with one annotated example. Decide how to treat shutters, partially hidden openings, temporary coverings and reflections before beginning the larger set.
Use photographs the team has permission to process. Preserve the originals and record image identifiers, capture position information available to the team and any prior resizing or cropping. If photographs include people or private interiors, address those permissions and the selected service's handling of the files before upload.
Review difficult edges deliberately
Begin with a small sample that includes clear views and harder cases. Compare the mask with the original at a useful inspection scale. Look for thin mullions, deep shadows, nearby vegetation and repeated openings that may be visually similar but belong to separate objects.
Keep the prompt and corrections associated with the accepted mask. A reviewer should be able to distinguish a model's initial prediction from the team's edited result. When the boundary remains ambiguous, retain an uncertainty label instead of forcing a clean silhouette for presentation.
For video, check frames where the object becomes occluded or changes apparent shape. A persistent identifier is useful only if it continues to refer to the same intended facade element. Do not count every frame-level mask as another physical object.
Separate selection from architectural conclusions
A mask records a selected image region. It does not by itself establish material composition, damage severity, physical area or compliance. If another process uses the region for those purposes, it needs its own evidence and review. In particular, pixel area should not be substituted for facade area without an appropriate geometric method.
Deliver the accepted masks beside the originals, the selection rule and a short exceptions register. For an office trial, record the time spent correcting masks and resolving uncertain boundaries. That effort is part of the workflow's cost.
The useful outcome is a consistently defined visual dataset that someone else can inspect. It can support later architectural analysis while keeping the source observation, the selected region and any subsequent interpretation clearly traceable.