My first gut reaction to this was: Have you tried using a vision-language model?
I used to have an elaborate Home Assistant setup with tons of zigbee sensors here and there telling me if front door is not closed properly, if the rice cooker is left on (via current monitoring), if the garage door is left open (with a magnetic contact sensor), if window is left open, if the soil is dry, etc. I had to keep changing batteries of all these things and manage flaky Zigbee connections in the past.
More recently, I've found the absolute most reliable, ~zero-maintainence solution to be an LLM. Home Assistant already has an LLM plugin, you can wire it up to a local LLM if you want, and you can just feed it an image and ask "Is this garage door open or closed" and I found that to be much better for my sanity than changing batteries every 6 months and dealing with flaky sensors sending me false alerts.
I even found it useful for "Did the cat vomit on the floor" (if so notify me and pause the cleaning robot schedule) and "Do any plants look like they need attention"
For TFA's use case, if you already have a security camera facing front, you can just feed it that image and ask "Are the garbage bins out"
Also, if you're worried about cameras sending home images to the cloud, there are non-cloud options. (I use Unifi cameras and Home Assistant gets a RTSP stream over LAN.)
So those tags are going to have to be a little cheaper before I’ll use them. :-)
Luckily these are only £5 each https://www.alibaba.com/x/1lBJtin?ck=pdp
The majority of cost is the wrover and the programming jig (which I think the manufacturer built to order, it’s very… home brew