Opendream: A layer-based UI for Stable Diffusion(github.com) |
Opendream: A layer-based UI for Stable Diffusion(github.com) |
https://github.com/Interpause/auto-sd-paint-ext https://github.com/thndrbrrr/gimp-stable-boy https://www.magicbrushai.com/
Also, "Normal" layered non-destructive operations are a couple of orders of magnitude faster and do not require 8Gb of VRAM per 512x512 patch, or work only with fixed set of buffer sizes, or any of other strange things SD comes with. Like, how a non-destructive controlnet layer would look in Gimp?
It's specifically why I've avoided diving too deep into "prompt engineering", because the kind of incantations required today just aren't going to be the way most people interact with this stuff for very long.
The difference between UIs is actually not very relevant today; by now the generic workflow for complex scenes is more or less obvious to anyone who spent time with SD.
- Draw basic composition guides. Use them with controlnets or any other generic guidance method to enforce the environment composition you want. Train your own controlnet if you need something specific. (lots of untapped potential here)
- Finetune the checkpoint on your reference pictures or use other style transfer methods to enforce the consistent style.
- Use manual brush masking, manually guided segmentation (ex. SAM), or prompted segmentation (ex ClipSEG) to select the parts to be replaced with other objects. The choice depends on your case and need to do it procedurally.
- Photobash and add detail to the elements of your scene using any composition methods you have (noisy latent composition, inpainting etc) with the masks you created in the previous step. Use advanced guidance (controlnets, t2i adapters etc)
- Don't bother with any prompts beyond very basic descriptions, as "prompt engineering" is slow and unreliable. Don't overwhelm the model by trying to fit lots of detail in one pass; use separate passes for separate objects or regions.
- Alternative 3D version: build a primitive 3D scene from basic props (shapes, rigs). Render the backdrop and separate objects into separate layers as guides. Use them with controlnets & co to render the scene in a guided manner, combining the objects by latent composition, inpainting, or any other means. This can be used for procedural scenes and animation (although current models lack temporal stability).
As long as your tool has all that in one place, it's a breeze, regardless of the UI paradigm (admittedly auto1111's overloaded gradio looks straight out of a trash compactor nowadays). I expect 2D/3D software integrations being the most successful in the future, as they already offer proven UIs and most desirable side features. The problem is that in the current state SD can't do much in the production setting, it's not a finished product - so there's not a lot of interest in software integrations just yet.
The pro tools that have incorporated generative AI into their workflows are not at all textual. The environment that popularizes this among the general public will look a lot more like canva or maybe Instagram than what's popular now.
I think there's so much unexplored potential in UI and workflows around generative AI, we've barely scratched the surface. Very exciting times ahead!
Thank you so much for sharing this. Civitai keeps bugging me to create an account. This doesn't seem to suffer from the same flaw.
You can train (finetune) your own on your reference material.
When a user does img-2-img on a layer does it use the context from other visible layers in the generation?
If the user generates a picture of a horse and rider to add onto another composition - they probably want to include the saddle.
Would be even more interesting to get an ANN middle system of ontology of the (finally) represented content in order to change the single items.
An internal representation of qualified structured items in space as part of the chain. Prompt > accessible internal representation > render.
I'd love a colab notebook if anyone has the skill and time to do so.
If anyone gets to making one before me, please leave a PR!
If you get to this before me, please create a PR!
It's all moot though, because as far as I know, there is no proper 2D graphics editing software that uses DAGs and nodes. Everyone just copies Photoshop. Especially Affinity, which is grating, given their recent focus on non-destructive editing. For some reason, node-based UIs ended up being a mainstay of VFX, 3D graphics, and VFX & gamedev support tooling. But general 2D graphics - photo editing, raster and vector creation? Nodes are surprisingly absent.
That's because non-destructive editing is mostly useful for animation, image series/sequences, and asset reuse, which are the most common in these fields. 2D artists have a different mental model, which is additionally set in stone by Photoshop and other software imitating it. Photographers use non-destructive editing, but mostly in simple cases because advanced things (retouching, creative compositing) can't and don't need to be done procedurally anyway.
I can see that being sensible for simple linear flows from one step to the next, with no branching merging, or connections that skip steps.
Seems to me that with any of those other things, a layered UI is going to start to break down a lot faster.
Stable Diffusion needs to go out to the masses to a greater degree. The unnecessary garbage complexity (eg Comfy's ridiculous noodlescape) that developers keep including into the UIs is holding Stable Diffusion back significantly from a greater mass adoption.
I wrote a typescript API generator for ComfyUI recently and having programmatic access to let you build and send the execution graphs is a game changer. Hoping to have time to release it soon. Same can easily be done for any other language. Exciting stuff!
There still are some issues with the eyes and a bit of flickering but at the speed everything is moving I wouldn't be surprised if this improves in a year or two.
Needless to say, there's still a lot of artistry involved in such a process so anything is yet to be completely automated.
[1] https://www.youtube.com/watch?v=tWZOEFvczzA
1. https://github.com/OpenDreamProject/OpenDream 2. https://opendream.ai/
If no real children were harmed to produce this stuff than it should be treated like any other extreme works of fiction (e.g. violence in video games, graphical descriptions in certain books).
Being disgusted is not grounds for banning something lol.
I’ve only just used Dall-E or SD with basic prompts, or sometimes using photoshop afterward. I’m curious what you’ve been able to come up with using your more complex pipeline.
There needs to be a REALLY FUCKING STRONG effort to kill all CP AI anything. Full stop.
AI should automagically report any attempt at CP.
I don't think an AI model could generate realistic CP without being trained on examples, which would mean there is literally no way for this assumption to be true.
Imagine getting reported because you generated an image of an anime girl deemed to be only 17.
I'd personally rather live in a world where people generate distasteful images with an AI and have that AI unconstrained than the inevitable one where everything gets locked down and run by large corporations who will ultimately create more harm than someone generating some lolicon.
A model by itself though… you might as well ask a pencil to report someone for drawing graffiti. It does not make sense.
Money laundering, CSAM, terrorism, drugs.
I for one would much rather give pedophiles an opportunity to fulfil their sexual desires through AI-generated pictures than real ones.
Of course, we can talk about the training material. Are there actual child porn images in there? I seriously doubt it but who knows?
And perhaps a case could be made that AI-generated child porn could be a gateway to invite people who then seek out non-generated material.
But I think these are separate discussions to be had.
So if either case applies, whether it's training based on certain images, or it becomes a gateway, these are discussions to be had directly relating to whether or not it should be classified as abuse material.
Additionally, I'm not sure if the recommended help methods by professionals who deal with pedophiles is to let them fulfill their specific fantasies without a care.
There are lots of really important discussions to be had, but they're all tied to each other basically. We can't separate them out, nor should we aim to.
Would you prefer an AI of "describe and print in 8k a body pit of objectional political dissidents thrown into a pit after being starved to death in a comical way such that my SV_BubbleTime can laugh at it as ooppsed to being offended by just how horrific this situation was HAHA"
The censorship isn’t OF those things. It’s in the NAME OF those things.
To me it's akin to encryption being used for illicit/illegal activities. Any tool that gives power to people can and will be abused by people you'd want nothing to do with.
What did you have in mind?
It all seems incredibly complex. Not a reason to "not try", but i suspect we'll struggle to implement even the most basic thing. And even then, take that basic thing and apply it to every software where users can input data.
Plus we'd have to convince everyone to do this. Automatic scanning and submission of data is not a well liked topic. Remember how Apple doing basic CSAM scanning was full of panic?
Even if a government _forced_ us to do this, jurisdiction alone would be a big question. Some serious questions that need some serious thought, imo. Is being hand-wavy even worth the time?
In my opinion instances should let the user decide what they find problematic or not and unless it's just spam they shouldn't ban instances.
I just wanted to clarify that I did not mean that these topics are all unrelated. When I wrote that these are separate discussions to be had, I was rather trying to imply that these questions are important enough to deserve an own treatment. However, I do agree with you that in the end, they all contribute to the question whether or not artificially generated child porn is abusive or not.
I do appreciate another sibling comment that points out the relation to other fictional child porn, such as literary works.
Additionally, I would like to add another dimension to the topic, namely that IMO, there is often an unspoken underlying assumption that portraits consumers of child porn as (potential) predators. However, unlike a juvenile delinquent who might find it cool to break into a local corner shop at night to steal some cans of beer, pedophiles are usually not attracted to child porn as a matter of choice. Like many sexual preferences, it is often innate, and can also be a burden to them: imagine you know that what gets you on is morally wrong, even a crime, and for most of your life you are forced to suppress your real sexuality as a consequence.
I'm thinking that fictional child porn, even when it's not AI-based but perhaps created with photoshop or in form of stories, could potentially help pedophiles to find ways of somehow dealing with their sexuality without actually preying on innocent children.
However, all of these thoughts come from a very naive understanding of the subject matter. Neither am I a pedophile myself, nor do I know anyone who is, nor am I a psychologist or something who works in the field. So I am very interested in corrections or additional options - especially, as I pointed out before, if they are done in equally civil ways as they have been so far in this thread.
So would it be okay to distribute "revenge porn" imagery after the subject has died?
[1]: https://en.wikipedia.org/wiki/Legal_status_of_fictional_porn...
If it's so easy i'd love to see your implementation that works multi-language, across all media types, for all jurisdictions and hell handles burden all the massive number of edge cases.
Or frankly, any implementation. Whatever you think is easy and everyone should be doing - please point to an E2E implementation of it. Maybe i misunderstand your scope. Something where if a user submits CSAM, or does something to some country authority..?
there’s also negative prompting, which tells the model don’t do these things.
hallucinations can still happen but it’s much easier than you’re making it out to be.
Please show me a model trained on adult humans being fine tuned to generate a child human (fully clothed).
SD-based models can generate child bodies just fine, based on lots of training with non-porn images.
> Please show me a model trained on adult humans being fine tuned to generate a child human (fully clothed).
That's going to be hard in the SD space since there are, AFAIK, no models trained exclusively on adult humans (not even the base models), so you'd have to scratch-train a new base model to do this. (A model fine-tuned on adult images is still going to have the influence of base-model training, and unless massively overfit—usually very much an anti-goal, the exception being age-filter models whose entire focus is controlling apparent age [0]—will not have lost much of the generalization capability of the base model.)
The generalization abilities of models are good enough in other contexts that its plausible that realistic nude children could be generated by a model with no nude children in its training data that was otherwise trained on both clothed children and nude adults. I have no plans on testing this, however.
[0] e.g., https://civitai.com/models/65214?modelVersionId=74332
I have not stated anything contrary.
> The generalization abilities of models are good enough in other contexts that its plausible that realistic nude children could be generated by a model with no nude children in its training data that was otherwise trained on both clothed children and nude adults. I have no plans on testing this, however.
Stable Diffusion has failed me for much simpler interpolations than the one you're describing. I don't believe this would work based on previous interpolations I have seen in Stable Diffusion. You can convince me by showing an example, but not by stating that it is possible without one.
And does my face compare to the ones in the dataset.
>Please show me a model trained on adult humans being fine tuned to generate a child human (fully clothed).
That's not something I can easily train for you ahaha
Yes, exactly! You're fine-tuning something with more data. How do you fine-tune adult porn models to create CP without CP data? If you do what you're saying you'll have adult bodies with child faces.
> That's not something I can easily train for you ahaha
Can you show me any example where such a thing has worked? I don't believe such an example exists.
I suspect we could use some sort of central management service for "Internet Reports". Ie to deal with jurisdictions, reporting something to the right people, etc, as well as the complexities involved with identifying people.
Either way i think you're underselling the complexity. Or maybe you think it's so easy but no one cares, /shrug. Seems a long list of questions i'd have before i could even begin to implement it.
It’s kinda like arguing that racism is a fact of life for some and we should allow them to vent their racism in a controlled environment. Hard disagree.
> Stable Diffusion has failed me for much simpler interpolations than the one you're describing
Its succeeded for me in much more complex ones. SD (even with the same exact set of checkpoint, LoRa, etc., and other workflow elements) isn't consistent across apparent-to-humans complexity levels in its generalization ability, and there are a very large combination of potential combinations.
> You can convince me by showing an example
Yeah, I'm not going to try to make simulated child porn for you, and while I would have sided with you in the debate over whether your earlier proof request amounted to a request for that, this one very clearly does.
Please, explain how I am "very clearly" asking for it. Please. Literally both sentences preceding the one you quoted (as in, the entire paragraph containing this sentence) are talking about interpolation. How do you jump from interpolation to CP??
You can use Stable Diffusion.
>If you do what you're saying you'll have adult bodies with child faces.
The body will be similar to petite people or similar to lolicon content traslated into "real life", it's not a far guess by the model.
The model do the same thing as with finetuning with my face, it understand that my images are the images of an adult male face and the knowledge will be added to the higher understanding of adult male and this include nudity, when you finetune on children it understand it's a human and it will added to the higher understanding of humans (And that includes nudity in various forms).
Maybe the model wouldn't be perfect with children, but it can't be with my nudes either, maybe I don't have nipples and the model doesn't know it; but the guesses that the model does are usually good enough to be considered realistic or plausible.
Yes, and as I've already told you, it doesn't work in Stable Diffusion, or at least I haven't seen any examples. Don't repeat this point again, show me an example.
> The body will be similar to petite people or similar to lolicon content traslated into "real life", it's not a far guess by the model.
Really? Show me an example. Don't just claim it, show me some place where this worked. You've mentioned fine-tuning a model with your face. Fine-tune a child model with your face and show me that it outputs something roughly adult-like. Or do the opposite and fine-tune an adult model with a child face and show me the model generates a small person. Show me that it's true, don't just claim it to be.