Ideogram 4 is HERE: The Ultimate JSON Prompting Masterclass! #384
FurkanGozukara
announced in
Tutorials
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Ideogram 4 is HERE: The Ultimate JSON Prompting Masterclass!
Full tutorial: https://www.youtube.com/watch?v=TW3MRdd0MV4
Learn how to run Ideogram 4 locally with SwarmUI and ComfyUI, download the required model bundle, use ready Turbo/Balanced/Highest Quality presets, and create accurate structured JSON prompts with Ultimate Image Captioner Pro for image recreation, text rendering, batch captioning, and training dataset preparation.
Tutorial Links:
🖼️ Ultimate Image Captioner Pro:
https://www.patreon.com/SECourses/posts/ultimate-image-captioner-pro-162527725
🐝 SwarmUI installer, model downloader and presets:
https://www.patreon.com/SECourses/posts/swarm-ui-installer-model-downloader-114517862
🧩 ComfyUI installer:
https://www.patreon.com/SECourses/posts/comfyui-installer-105023709
🎬 Windows requirements tutorial:
https://youtu.be/DrhUHnYfwC0
⚙️ Requirements post with links/screenshots:
https://www.patreon.com/SECourses/posts/requirements-tutorial-step-by-step-written-111553210
💬 Discord support:
https://discord.com/invite/software-engineering-courses-secourses-772774097734074388
⭐ SECourses GitHub:
https://github.com/FurkanGozukara/Stable-Diffusion
Chapters:
00:00:00 Ideogram 4 overview: JSON prompting, SwarmUI presets, ComfyUI workflows, and model bundle
00:00:53 Ultimate Image Captioner Pro for turning reference images into Ideogram JSON prompts
00:01:10 Editing JSON elements, bounding boxes, wanted text fields, captions, and prompt layout
00:02:02 Regeneration examples showing structure, objects, scene layout, and image text matching
00:03:18 Captioner Pro feature tour: Qwen, JoyCaption, saved outputs, and JSON builder
00:04:30 Dataset workflow: prompt presets, batch folder captioning, and automatic VRAM presets
00:05:13 Tutorial roadmap: ComfyUI update, SwarmUI update, model download, app install, usage
00:05:40 Updating ComfyUI by extracting the latest installer zip and overwriting old files
00:05:56 Optional fresh ComfyUI venv rebuild for fixing outdated or broken installations
00:06:15 Running the ComfyUI update script, Python choice, UV speed, and quant support
00:07:05 Installing recommended custom nodes bundle 100 for ComfyUI and SwarmUI compatibility
00:07:50 Launching fresh ComfyUI and testing the Ideogram Turbo preset workflow
00:08:47 Setting width, height, resolution, and matching prompt aspect ratio
00:09:07 Updating SwarmUI with the latest zip, overwrite method, and safe folder paths
00:09:48 Automatic .NET SDK 10 install and why SwarmUI needs the correct SDK version
00:10:51 SwarmUI backend setup: ComfyUI backend, Triton, Sage Attention cautions, extra args
00:11:44 Downloading the Ideogram 4 core bundle with hash verification
00:12:28 16-connection parallel downloads, target folders, ComfyUI mode, and URL downloader
00:13:20 Merging model parts and sharing SwarmUI models through extra_model_paths.yaml
00:13:51 Setting SwarmUI model root to reuse another model folder and avoid duplicates
00:14:12 Updating SwarmUI presets with delete import, normal import, overwrite, and backup
00:14:58 Refreshing presets and confirming Ideogram Turbo, Balanced, and Highest Quality
00:15:14 First simple Ideogram prompt, false safety filter block, and weak plain prompting
00:15:34 Using Realism Engine Ideogram 5 LoRA to fix the blocked car prompt
00:15:57 Why detailed JSON prompts are needed and downloading Captioner Pro
00:16:23 Installing Captioner Pro with Windows install update app, venv, and model downloads
00:16:34 Windows requirements: Python, CUDA, cuDNN, C++ tools, FFmpeg, Git, and setup guide
00:17:03 Cloud/Linux notes plus Massed Compute interface, creator image, GPU, and coupon
00:17:34 Captioner installer downloader: 16 connections, hash checks, and accurate setup
00:17:57 Starting Ultimate Image Captioner Pro and saving custom user presets
00:18:14 Loading the Bugatti reference image and generating official Ideogram JSON
00:18:39 Prompt generation speed, copying the prompt, and understanding VRAM usage
00:19:09 Subprocess mode to release all VRAM and RAM after each captioning run
00:19:54 Reviewing generated JSON: high level description, visible text, boxes, and details
00:20:21 Pasting JSON into SwarmUI and matching the custom 5:3 aspect ratio
00:20:43 Aspect ratio calculator, side length control, and high resolution generation
00:21:36 Comparing with and without aspect ratio metadata and avoiding false safety blocks
00:21:58 Realism Engine LoRA strength, when to use it, and output comparison
00:22:34 Choosing Turbo, Balanced, or Highest Quality and testing Turbo speed
00:22:54 Ideogram 4 image to image, inpainting, image creativity, and image prompts
00:23:19 Captioner Pro batch folder processing: subfolders, overwrite, and append modes
00:23:35 Post processing captions with prefixes, suffixes, replacements, and sensitivity
00:24:07 Final options, auto quantization by GPU VRAM, support channels, and closing
Covered in this video: local Ideogram 4 installation, SwarmUI and ComfyUI preset usage, automatic model downloads, JSON prompt creation, bounding box editing, image recreation, safety filter fixes, LoRA realism settings, batch captioning, and VRAM friendly caption generation.
Video Transcription
00:00:00 Greetings everyone, today I am going to show you everything about newest Ideogram 4 model. This
00:00:06 model is extremely powerful, it works and supports with JSON prompts, and how to use it is not easy
00:00:14 as other models, but it is very powerful. I have prepared 3 different presets, they are all ready:
00:00:21 balanced preset, highest quality preset, and turbo preset. These are SwarmUI presets but
00:00:27 their same versions exist on ComfyUI as well as workflows. Downloading the necessary models is
00:00:34 also ready with our model downloader Ideogram 4 core bundle, you can download them right away or
00:00:41 you can go to the image generation models and see the additional models here as well. So it
00:00:47 works with JSON prompts but how to use this model with JSON prompts easily? To generate
00:00:53 JSON prompts I have developed a new application called as Ultimate Image Captioner Pro. This
00:00:59 application is a powerhouse, I will show all features of it. You see this is the input image,
00:01:05 this is the JSON prompt it generated, these are the JSON elements that you can modify, and this
00:01:10 is the visualization of the JSON prompt. Another example, this is input image, this is JSON prompt,
00:01:15 you see these are the captions of the boxes, when we scroll down we can see the boxes here,
00:01:21 you can move these boxes very easily like this. If you want to work with the boxes more easier,
00:01:26 hide them, and drag and drop, make the changes like this, and once you done with everything,
00:01:32 including changing the captions or everything, apply box edits and it will save the edited
00:01:38 boxes. So you see as you click here and here it will update all these values for you, you
00:01:44 can change them. One of the big advantage of this model is that it has both caption and text area.
00:01:50 So you see it is taking the text differently. You see this is wanted text and the wanted text
00:01:55 is written like this. Caption is the regular caption, text is if there is a text. The model
00:02:01 works without any JSON prompt as well, it is not mandatory but the power of the model comes from
00:02:07 JSON. So this is another example as you see, and I have used these examples and regenerated images
00:02:13 like this one. So this and this image is perfectly matching as a structure, the style is different,
00:02:19 obviously you need to change the style, but you see there is a moon here and we can see that in
00:02:25 the prompt there is a crescent moon visible in the dark sky. So it is perfectly matching. There are
00:02:30 2 mouses, we can see how it is captioned, it says a white mouse wearing a beach coat and blue scarf
00:02:37 holding a snowball, smiling with snow on its head. So I need to change the style if I want to change
00:02:42 the style, there is no style information here, but the structure is fully matching. Mouse colors,
00:02:48 what they are wearing, what they are doing, the overall scene. This is another example, this is
00:02:53 generated from this, you see it was like this and this is the generated example. So the model is
00:02:59 perfectly able to regenerate the original input image, as you wish. So this is the regeneration
00:03:05 of this original image. You see it was this one and this is regeneration. So this model is very
00:03:12 powerful as I said, you can use it as you wish. With this captioner application you won't have any
00:03:18 issues to regenerate any images. This application has so many features, not only Qwen Instruct,
00:03:24 but we also support JoyCaption as well. So you can use JoyCaption captioners and you will see
00:03:31 the generated captions and everything. Everything is automatically saved inside outputs folder.
00:03:37 Moreover, we have JSON builder as well. So with JSON prompt builder you can start from beginning
00:03:42 or you can load your existing generations and make changes and save everything. This is amazing. The
00:03:49 JSON prompt builder is so easy to use, when you click generate JSON it will start empty like this,
00:03:54 so you can keep adding boxes, type captions, whatever you want, and save them. So it will
00:04:00 be saved in a new fresh folder, but I prefer to load the generated JSON and work on it. It is up
00:04:07 to you, you can do both ways. Finally you can also use the saved outputs. So with saved outputs you
00:04:14 can refresh and when you click it will load very easily all the generations. There is filtering,
00:04:20 daily filtering, number of displayed results per page, it is very advanced, you can use this very
00:04:26 easily. So if you are going to make training, this application is perfect for that. You see we
00:04:30 support all these prompting presets: Ideogram, photorealistic, art style, character subject,
00:04:36 product object. So you can use these according to your dataset and you can do batch captioning. In
00:04:42 the bottom we have batch folder captioning, it will use these set parameters and batch caption
00:04:48 the given folder and save all the outputs. Moreover, we support all these VRAMs. So
00:04:53 if you have a low VRAM GPU, don't worry we cover you. These low VRAM presets exist for JoyCaption
00:05:00 as well. So whether you are using JoyCaption or Qwen, you can select your VRAM preset. It will
00:05:06 automatically select it according to your GPU, but you can go to the lower or higher VRAM presets as
00:05:13 well. How you are going to use this model? I will begin with showing you how to update your ComfyUI,
00:05:20 then how to update your SwarmUI, then how to download models, then how to install the image
00:05:26 captioner app, and the rest is as usual using. All older tutorials are valid, there aren't many
00:05:33 new stuff, so let's begin. So first of all, check below description and go to ComfyUI installer and
00:05:40 download the latest version. Move the downloaded zip file into your ComfyUI installation folder.
00:05:46 Right click and extract and overwrite all the files. This is super important. Once you
00:05:51 overwritten all files, you are ready to update. I recommend you to delete your virtual environment
00:05:56 folder if you didn't update for a long time, so that we will get all freshly installed virtual
00:06:02 environment. This is not mandatory step, but this is perfect way of fixing your ComfyUI installation
00:06:08 and updating it. Okay, the folder is deleted. Then run the Windows install or update ComfyUI.bat
00:06:15 file. Choose your Python version, currently I am still using 3.10 but we may move to 3.11 as
00:06:22 a preferred Python version soon because a lot of packages are moving to that. This will generate
00:06:27 the virtual environment or update the libraries if it is needed and it will update your ComfyUI to
00:06:32 the latest version. Moreover, it will update the some of the custom nodes that we use by default
00:06:38 with our workflows and presets. We are using UV installer, therefore the updates or installations
00:06:45 are super fast. Moreover, I am automatically installing quant operations, therefore our
00:06:51 ComfyUI automatically supports all of the quants versions that you might find out there, like int8
00:06:58 block-based quantizations, different quantizations there you might find. Okay, it is all done,
00:07:05 everything is set. One more thing that I recommend you to do is use the Windows custom nodes bundles
00:07:12 installer, run, and my recommended bundle is 100. You can also choose them 1 by 1 from here
00:07:19 with comma separation, but I am going to choose bundle 100 and hit yes. This will update the most
00:07:26 commonly used nodes to the latest versions. You see these nodes will get installed, this is what I
00:07:33 recommend to use with ComfyUI and SwarmUI with the maximum quality and performance. So it is updating
00:07:39 my nodes and everything should be ready. If you use other different custom nodes, then these ones,
00:07:45 they may conflict with your installation, they may break your installation, but these are the my
00:07:50 recommendation. So my ComfyUI is now ready. I am using the extra model paths.yaml file, therefore
00:07:58 it is seeing all of the models downloaded into my SwarmUI. So now I can start and use right
00:08:04 away. Actually let me show you, run GPU.bat file, it will start the ComfyUI freshly set,
00:08:09 everything is freshly set right now. We are still using PyTorch 2.9.1 but I plan to upgrade
00:08:16 to PyTorch 2.12 soon, once the TorchAudio is also updated. Then inside the presets you will see the
00:08:24 Ideogram presets. You can use any of them. Turbo is also working very well. When I drag and drop it
00:08:30 will be loaded like this. You see all the models are automatically seen accurately. When I run it,
00:08:35 it should pretty fast generate output. The turbo preset is really fast. And it is done. The turbo
00:08:41 preset already generated. One more thing that I need to mention is that you need to set the
00:08:47 width and height from here. These width and heights are not important, so set your width
00:08:53 and height from here to set your resolution. Moreover, the prompt generator uses aspect ratio,
00:09:00 therefore try to match your aspect ratio with your resolution to the prompt, or change both of them
00:09:07 accordingly. Now as a next step we will update our SwarmUI. So go to the description below and go to
00:09:14 the SwarmUI link and download the SwarmUI model downloader zip file. Move it wherever you are
00:09:20 going to install or your existing installation. So I recommend you to not have any special
00:09:27 characters in your folder paths, including spaces or non-English base characters. Make your folder
00:09:31 paths like this. Then right click and extract and overwrite all the files. This is super important,
00:09:37 overwriting all the files. Once it is done, you can just use install SwarmUI or update SwarmUI.
00:09:43 There is 1 more thing that I want to mention before I update it. When you start the SwarmUI,
00:09:48 you may have noticed that it is telling you this: please install .NET SDK 10. So the SwarmUI
00:09:55 is going to update to SDK 10 version. I updated our installer and updater, when I update SwarmUI,
00:10:02 now it will automatically install the accurate SDK version. It is also going to ask permission,
00:10:09 that is why I deleted it from my computer to show you. Okay, now it is asking the permission,
00:10:14 I click yes, and it will open this screen and install. It will first download the exe file,
00:10:19 then it will start the installation process like this, that you need to click and continue. This
00:10:24 way you will get the accurate .NET SDK version and you will have the latest version. So your
00:10:30 SwarmUI will keep working. This is necessary since SwarmUI is actually programmed with C# rather than
00:10:38 Python. We are using ComfyUI as a backend if you remember my previous tutorials. So the backend
00:10:43 is ComfyUI but SwarmUI is basically a wrapper that lets you use the ComfyUI with much easiness. Okay,
00:10:51 it is done. Then it will continue updating, it will update the necessary other stuff if there is
00:10:56 anything, it will compile and start the SwarmUI. Once SwarmUI started, make sure that you are using
00:11:02 our ComfyUI backend installation, and now I am using enable Triton backend. I don't recommend
00:11:08 you to add Sage attention by default, because in some models, in newer models, it may not work very
00:11:14 well. So you need to test whether it is working or not on each model that you are using. Enable
00:11:20 Triton backend is working amazing. Enable Triton backend is also added to the ComfyUI starter,
00:11:26 when you edit the run.bat file you will see that it is using the enable Triton backend.
00:11:32 So it also uses Sage attention, so you can remove it from there as well if you need. This is how you
00:11:37 add extra arguments to your ComfyUI backend from SwarmUI interface. So as a next step you need to
00:11:44 download necessary models. To download necessary models we are going to use start download models
00:11:50 app.bat file. It will start the model downloader application. And then in the SwarmUI bundles you
00:11:56 will see that Ideogram 4 core bundle. Download all the models, it will download if they are missing,
00:12:02 if they are already downloaded it will just hash verify them. SwarmUI may modify your downloaded
00:12:09 models and it will cause mismatch of the hash files. In that case it will redownload. But this
00:12:16 downloader is made very well, it verifies hash files, so with this downloader you will never have
00:12:22 corrupted model issues. Moreover, it starts 16 different parallel downloads, therefore it is able
00:12:29 to download with maximum speed that your internet service provider supports. Currently you see it is
00:12:34 downloading with 100 megabytes per second, this is my maximum speed, I have 1 gigabits internet.
00:12:40 And we can also see the downloaded model parts here, so you see it is downloading as a 16 parts,
00:12:47 16 connections. All the models, this application download, downloaded same way. You can set your
00:12:52 target model folder, it also supports ComfyUI model structure, just enable this checkbox. It
00:12:58 also supports Forge WebUI Automatic1111 folder structure or lower case folder names. Moreover,
00:13:04 it also has URL downloader, I had explained all of this in previous tutorials, so you can download
00:13:09 from Civitai or Hugging Face into target folder. And it will be very fast and hash verified. I
00:13:15 really recommend to use this model downloader. Once the model downloaded, they will be merged
00:13:20 into single part. And I am not duplicating the models, I am using the extra model paths.yaml, you
00:13:27 need to copy this and paste it into your ComfyUI folder. It is coming with our zip file. And when
00:13:32 you edit this file you will see your base folder path, you need to change this according to your
00:13:39 SwarmUI installation, therefore it will see all the models that was downloaded into your SwarmUI,
00:13:45 so that you can use it inside ComfyUI as well. Or in the SwarmUI, I think it also supports that,
00:13:51 so go to server configuration and you see there is model root, so you can give another root folder
00:13:57 like your ComfyUI models, and when you save it in the bottom, or auto save it, yes it is probably
00:14:02 auto saved as you change them, so it will see the models from that another folder as well. You don't
00:14:07 need to duplicate any models. Then what you need to do is, you need to update presets. For updating
00:14:13 presets I recommend you to use Windows preset delete import. It will clear all of your presets
00:14:20 and update them to our latest versions. You can alternatively also use import, so choose file,
00:14:26 go to the folder and pick the amazing SwarmUI presets, currently version 51. It will say that
00:14:33 are you want to overwrite or not, then you can overwrite and import everything. Alternatively,
00:14:38 you can click the Windows preset delete import, you need to run this when the SwarmUI is running,
00:14:45 then it will ask you whether you are sure or not, yes, and it will clear all of your presets
00:14:50 and import them like this. It also backups your presets inside utilities folder as presets backup
00:14:57 before deleting them. So once the presets are refreshed you will see them like this, so refresh
00:15:02 presets, and the Ideogram presets have arrived. Now they are ready to use. Once the models are
00:15:08 downloaded, yes I see they are already downloaded and verified. So for using as always, as usual
00:15:14 quick tools, reset params to default, then select the preset that you want to download, like let's
00:15:19 select turbo direct apply. And you can type your prompt. Amazing car going fast on a road. This
00:15:27 is a very simple prompt. And unfortunately it is blocked. So we have a LoRA, Realism Engine
00:15:34 Ideogram 5, it is also downloaded with the core bundle, let's select this and try again. Let's
00:15:39 see if it will fix this issue. Yes, you see the previous image was blocked by the safety filter,
00:15:45 but when I enabled the Realism Engine Ideogram version 5, it is working. However, it is not a
00:15:51 good quality because this model wants you to have detailed JSON prompts. So I am going to do that,
00:15:57 how? Check below and you will see Ultimate Image Captioner Pro link, open the page and download
00:16:04 the Ultimate Image Captioner Pro latest zip file. I will show a fresh installation into my Q drive,
00:16:11 so I will paste it there, right click and I will extract. Then enter inside the extracted folder
00:16:16 and all you need to do is just use the Windows install update app.bat file. It will install,
00:16:23 it will install the application with a virtual environment and download all the necessary
00:16:27 models. This application is using Python 3.11. Always pay attention to the Windows requirements,
00:16:34 if you didn't watch the requirements tutorial previously please watch it, its link is fully up
00:16:39 to date, so when you open the link of the tutorial you will see everything with images as you are
00:16:45 seeing right now. So you won't have any issues how to follow tutorial, how to use the latest
00:16:51 updated libraries. When you follow this tutorial you will be fully ready to run any AI application
00:16:57 on your Windows computer. As always we also have RunPod and Massed Compute instructions, so if you
00:17:03 are a Linux user you can use the Massed Compute instructions. Everything is fully up to date,
00:17:09 even the tutorial videos are up to date in the instructions read.txt file. There is only 1 thing
00:17:14 that I want to show you, Massed Compute updated its interface, so if you use Massed Compute make
00:17:20 sure that category creator, image SECourses, select your GPU, enter your coupon as SECourses
00:17:28 and verify. You see you will get amazing discount in all of the GPUs that you can use on Massed
00:17:34 Compute. So the application is getting installed and it is downloading the models, I already have
00:17:40 it in somewhere else. This is also using 16 connection download and hash verification. All
00:17:46 of my installers uses this specific special downloader that I have developed, therefore
00:17:52 they will be always fully accurate. So the other application was installed here, I will just run
00:17:57 it with Windows start Ultimate Image Captioner Pro. Once your installation has been completed,
00:18:02 this is the interface. It supports custom user presets as well, so you can make changes and save
00:18:08 them. So let's find an image and replicate it in the Ideogram 4 model. For example, let's try
00:18:14 this image. I will pick the image, so for picking image click here or you can drag and drop. Okay,
00:18:20 this is the image. Then I will caption image, I am using the Ideogram official version 1 preset,
00:18:26 you can also use other presets as I have said. It also supports text generation, not only JSON
00:18:32 based prompts but also text based prompts as well. Moreover, our application is super optimized, it
00:18:38 will also even show you the token speed, let's see the token speed, it is about 25 token per second.
00:18:45 It is up to 4000 tokens, but usually it ends much faster depending on the image. So it is generated
00:18:51 in 11 seconds. Let's copy this prompt. By the way, now it is keeping my VRAM busy, how? When I open
00:18:59 it I can see that if you don't want to keep your VRAM busy, what you can do is, let me close this
00:19:05 and show you again. So other applications are also keeping VRAM busy. Let's terminate them and let's
00:19:11 start the image captioner Pro again. This way you can run both of the applications at the same
00:19:15 time. Okay, so our VRAM is like this right now. This is SwarmUI started the ComfyUI backend. So
00:19:22 for Ultimate Image Captioner to not use any VRAM, there is this option: run single and batch in sub
00:19:29 process. So what will this do? This will generate caption and then terminate the process. Therefore,
00:19:35 it will leave absolutely 0 VRAM and RAM usage. It will fully terminate process. This way you
00:19:41 can keep your captioner open, run it anytime you want, and it will fully terminated later. It won't
00:19:48 have any VRAM usage, any VRAM leakage, we will see in a moment, yes, it terminated and generated the
00:19:54 caption. So let's verify that it is accurate, yes, we can see that a black, orange Bugatti Chiron,
00:20:00 and it even shows its text here. You see, amazing. We can see also generated caption. So there is 1
00:20:08 high level description, it says a black and orange Bugatti Chiron Super Sport 300 plus sports car.
00:20:16 This is the main structure of the image, then it shows the text as well. So let's return back to
00:20:21 our SwarmUI. And remember last time it had given us image blocked by safety filter, just paste
00:20:29 this. Let's see the aspect ratio to be sure, okay, it is 5 to 3. So this is a pretty custom aspect
00:20:36 ratio, perhaps we can, yeah, there is not exactly 5 to 3. So how we gonna do that 5 to 3? I asked
00:20:43 the SwarmUI developer to add custom aspect ratio here, I hope he adds, you can also tell him. There
00:20:50 is a website that I have found, aspect ratio calculator, so let's enter here, let's set our
00:20:56 aspect ratio, and let's generate a high resolution image. So it is around, yeah, 2045 to this one,
00:21:04 so I will enter custom, this will be this, and this will be that. If it was a standard aspect
00:21:10 ratio it would be much easier, or alternatively we can remove the aspect ratio from here, but it may
00:21:17 break our, yes, B boxes probably, let's try it. So let's also select, then you can have side length,
00:21:24 so this is very useful, why? This way I can change the resolution of the image while keeping the
00:21:30 aspect ratio, this is a new feature and it is amazing. Okay, we got our image, pretty cool.
00:21:35 Let's try this one and see without aspect ratio how it works. Okay, it is I think generating
00:21:41 almost same image. This image was, let's see the generation, yeah, 2045 to 1227, this will be 2016
00:21:50 to 1152, yeah, and this is another image. So when you do JSON prompting, the stupid safety filter
00:21:58 should not be triggered unnecessarily. And in the LoRAs you can set your Realism Engine Ideogram 5,
00:22:04 this is automatically downloaded, if you want to change its impact strength, you can change it
00:22:09 from here, you see, with my mouse, or you can just type it, like let's try 1.1. And this is improving
00:22:16 the realism, I tried it, it is pretty useful. So if your prompt is related to realism, use this,
00:22:22 if it is not, then don't use it. And it is ready. So with realism I got this, without realism I got
00:22:28 this. It is up to you, you can try them. You can change the presets, this was the turbo preset,
00:22:34 there is also balanced preset and highest quality preset, you can use any of them,
00:22:38 but turbo is pretty fast and especially fast at 1024 pixel, let's see real time how fast it is
00:22:45 being generated. Okay, almost ready, yes. So it took like, let's see, 7 or 8 seconds to generate.
00:22:54 The model supports image to image as well, so you can use init image, set image creativity, and use
00:23:00 the prompt from this to this, so you can do image to image as well, it is pretty consistent, it also
00:23:06 supports inpainting. It also even supports image prompt from here, upload prompt image, however I
00:23:13 didn't test it and it is not very good like Qwen 2512, so you need to play with it, you can provide
00:23:19 image prompts as well. And about image captioner app, for folder batch processing enter your
00:23:25 folder, input folder here and output here, you can also process sub folders if you want, you can also
00:23:30 overwrite captions or append captions, it supports all of them. And 1 another thing is that we
00:23:35 support text prefix and suffixes, so you can add OHWX to all of your captions automatically, or you
00:23:42 can even have replace word, like man replace it with OHWX and add, and it will add it as a replace
00:23:49 word. And it will post process all generated captions and replace these words. You can make
00:23:54 it case sensitive or single word sensitive. When you make it single word sensitive actually it will
00:24:00 only find man word and not the wildcard prefix. So it is all up to you, we have all the options.
00:24:07 You can also disable auto save box image. So check all the options we have, these are all set after
00:24:15 deep research and everything is automatically set. You see when I change the preset it changes the
00:24:20 quantization fully automatically, everything is fully working automatic depending on your GPU. So
00:24:25 you see 16 gigabytes preset using int8 version, 24 uses BF16 highest quality, 10 gigabyte uses NF4.
00:24:34 So this application is really good. You can always ask me any questions that you have from Discord
00:24:39 or from YouTube as a comment or from Patreon. I hope you have enjoyed, hopefully see you later.
All reactions