Lots of possibilities for expanding the structure of a scene. You can repeat categories, like “Person-Action-Object-Person.” Or you can add categories like “Person-Action-Adjective-Object.” You could go crazy and add multiple new categories and repeat them like “person-action-object-person-adjective-animal-food-vehicle” and cram 16 digits into a scene.
Keep in mind that you aren’t gaining any better data compression by adding more elements to a scene. You still have to encode and decode an intentionally created element for every 2 digits you want to memorize. Usually less complex scenes are easier to encode and recall.
The “advantage” that some will point to is that you need less loci to encode the same number of digits, like if you want to memorize 24 digits, you need to use four loci with 6-digit scenes and only three loci if you create 8-digit scenes, but usually it is MUCH MUCH easier to just use a couple more loci to store the digits than it is to build and accurately recall the more complicated scenes. A huge benefit to using memory palaces is to take advantage of the ease of use and simplicity it offers by giving you the ability to spread information out across instinctively memorable locations. You lose a lot of this benefit if you try to cram too much into a single location.
I build out a Person-Action-Adjective-Object system (fully based on Major) a while ago and used it successfully to memorize thousands of Pi digits. I’ve since moved on a single-association 3-digit system and haven’t regretted the change one bit, as the simplicity in recognition, scene building, and recall is much improved!
I honestly recommend most people DROP an element and go with a simplified PO or PA structure for a 2-digit system as it makes scene generation much more natural and usually faster. Or make the jump to a 3-digit system. I’d recommend that before I’d suggest adding additional categories.