We parse the input and output scenes into objects.
For each input object, there is only one object in the output scene that it is contained within, making it straightforward to produce a mapping from input objects to output objects.
e.g.
All object mappings involve replacing the input object’s shape with a common output shape, making it straightforward to produce the single required rule:



