FLUX 2 vs FLUX SRPO, New FLUX Training Kohya SS GUI Premium App With Presets & Features #349
FurkanGozukara
announced in
Tutorials
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
FLUX 2 vs FLUX SRPO, New FLUX Training Kohya SS GUI Premium App With Presets & Features
Full tutorial: https://www.youtube.com/watch?v=RQHmyJVOHXo
FLUX 2 has been published and I have compared it to the very best FLUX base model known as FLUX SRPO. Moreover, we have updated our FLUX Training APP and presets to the next level. Massive speed up gaings with 0 quality loss and lots of new features. I will show all of the new features we have with new SECourses Kohya SS GUI Premium app and compare FLUX SRPO trained model results with FLUX 2.
Get the SECourses Premium Kohya Trainer DreamBooth / Fine Tuning : [ https://www.patreon.com/posts/Kohya-FLUX-DreamBooth-Trainer-App-112099700 ]
Get the SECourses Premium Kohya Trainer LoRA : [ https://www.patreon.com/posts/Kohya-FLUX-LoRA-Trainer-App-110879657 ]
DreamBooth Training Tutorial: [ https://www.youtube.com/watch?v=FvpWy1x5etM ]
LoRA Training Tutorial: [ https://www.youtube.com/watch?v=nySGu12Y05k ]
Qwen Image Realism Tutorial: [ https://youtu.be/XWzZ2wnzNuQ ]
Join our Discord Community: [ https://discord.com/servers/secourses-Discord-772774097734074388 ]
⏱️ Video Chapters:
00:00:00 Introduction to New FLUX Training Improvements and Local Training Showcase
00:00:24 Understanding FLUX SRPO Model: High Realism with Minimal VRAM Requirements
00:00:38 Updated Configurations for Training Realism on 6GB VRAM GPUs Locally
00:01:07 FLUX 2 Announcement and Setting Up Comparisons with BFL Playground
00:01:45 FLUX 2 Dev Model Technical Specs: 32 Billion Parameters and Hardware Challenges
00:02:11 Overview of Changes in SECourses Premium Kohya Trainer Version 35
00:02:46 Development Updates: GUI Improvements and Full Torch Compile Support
00:03:13 LoRA Presets Update: VRAM Optimization and Speed Improvements via Torch Compile
00:03:27 Introducing On-the-Fly FP8 Scaled LoRA Training Support
00:03:42 Quality Comparison Analysis: BF16 vs FP8 Scaled Weights LoRA
00:04:24 VRAM Usage and Speed Analysis: Block Swap Count Reduction with FP8 Scaled
00:05:12 Why 32GB GPUs Don't Need FP8 Scaled: Fitting Completely with Torch Compile
00:05:36 DreamBooth Training Specifics: BF16 Mixed Precision and Tab Selection
00:06:18 New 80GB GPU Configuration: Significant Training Speed Up and Cost Analysis
00:07:39 New Feature: FLUX FP8 Converter Tool for DreamBooth Models
00:08:10 Verifying Quality of Converted FP8 Scaled Models vs Original BF16
00:09:03 New Image Pre-processing Tool: Visualizing How Kohya Sees Your Dataset
00:09:34 Demonstration of Pre-processing: Identifying Padding and Orientation Issues
00:10:38 Additional Features: Memory Efficient Loading and CPU Text Encoder Caching
00:11:23 New Automated Model Downloader Tool: Installation and Model Selection Guide
00:12:35 FLUX 2 vs FLUX 1 Context: Can the New Model Replace Existing Workflows?
00:13:03 Setting Up the Generation Comparison: SRPO Fine-tune vs Base vs FLUX 2 Pro
00:14:13 Analyzing First Comparison Results: Portrait Quality and Realism Assessment
00:15:02 Second Comparison Test: Handling Unrealistic Elements and Animals in Prompts
00:15:33 Importance of Resolution and Platform Choice for FLUX 2 Quality
00:16:24 Analyzing Second Results: Prompt Adherence and Realism in FLUX 2 vs SRPO
00:17:07 Testing the Prompt on Nano Banana Pro: Quality Assessment
00:17:39 Comparison with Seedream 4 Model and Final Thoughts on FLUX 2 Potential
00:18:50 Current Recommendation: Why Qwen Image Realism is the Temporary King
In this video, I demonstrate the newest improvements and features added to our SECourses Premium Kohya Trainer (v35). We have achieved massive performance gains and VRAM reductions, allowing for high-quality FLUX training on GPUs with as little as 6GB of VRAM using the FLUX SRPO model.
I also break down the brand new FLUX 2 announcement! We perform a side-by-side comparison between my locally trained FLUX SRPO model, the base FLUX 1 Dev, and the newly released FLUX 2 Pro model to see if it’s time to switch workflows.
🚀 Key Updates Covered:
Full Torch Compile Support: Now enabled across all presets for faster training speeds.
FP8 Scaled LoRA Training: On-the-fly conversion that drastically reduces VRAM usage (0 block swaps on 24GB cards) with zero quality loss.
New Tools: Introducing the "FLUX FP8 Converter" to shrink DreamBooth models to 11.1GB and a new "Image Pre-processing" tool to visualize exactly how Kohya sees your dataset.
Hardware Optimization: New presets ranging from 6GB VRAM consumer cards up to 80GB A100/H100 configurations for lightning-fast training.
FLUX 2 Analysis: A realistic look at the 32-billion parameter giant and how it compares to our optimized FLUX 1 workflow.
Video Transcription
00:00:00 Greetings everyone. Today, I am going to show you our newest improvements and features regarding
00:00:07 FLUX training. These images that I am showing you right now were trained locally with our
00:00:14 newest FLUX Kohya trainer, improved by SECourses (by me), and I have used the FLUX SRPO model.
00:00:24 The FLUX SRPO model is the same as the FLUX Dev model; it uses a minimal amount
00:00:30 of VRAM and doesn't require strong hardware, but it is a very realistic model. With our updated
00:00:38 configurations and training presets, you will be able to train amazing realism-having FLUX
00:00:45 SRPO models, the same as the FLUX Dev model. Our presets are working both on the FLUX Dev
00:00:52 model and the FLUX SRPO model, and as I said, it works as low as with 6 GB of VRAM-having GPUs;
00:01:01 you can train locally on Windows. Furthermore, FLUX 2 has been published
00:01:07 today, a few hours ago, and I am going to make a comparison with FLUX 2. I am going to
00:01:14 use the BFL (Black Forest Labs) playground to make comparisons. For example, this image was generated
00:01:22 with FLUX 2 at 2048 by 2048 pixels. And this image was generated locally on my computer with the FLUX
00:01:31 SRPO model using our FLUX generation preset, and the biggest difference is that the FLUX SRPO model
00:01:37 is very small compared to the FLUX 2 Dev model. The FLUX 2 Dev model is massive; it is 32 billion
00:01:45 parameters, 64.4 GB, and this is BF16, not even FP8 quantized. ComfyUI is already working on
00:01:56 quantized models for the FLUX 2 model. You see, its text encoder is 35 GB. Its quantized FP8 model
00:02:04 is 35 GB. Therefore, currently running the FLUX 2 model on our consumer GPUs will be very hard.
00:02:11 So, what changes do we have with our Kohya trainer, the SECourses Premium Kohya Trainer?
00:02:18 Let me show you them quickly. So, our training tutorials are still 100% valid. You see the
00:02:25 link is here. This is DreamBooth training; this tutorial includes Windows training,
00:02:30 RunPod training, and MassedCompute training. And if you are interested in LoRA training,
00:02:34 this tutorial has LoRA training for Windows; we also have RunPod and MassedCompute as well.
00:02:40 The links will be in the description of the video. So, what changes have we made since these
00:02:46 tutorials? With version 35, we have added so many new features. Now I am developing the GUI
00:02:53 and the SD Scripts; I have forked their original repositories and I am improving them. We are fully
00:02:59 supporting Torch Compile. Let me open all sections to show you the newest changes. So when you search
00:03:06 for "Torch," you will see Torch Compile; all of our new presets are supporting and using this.
00:03:13 The new presets are like this. You see, these are our LoRA presets. I have updated all the VRAM
00:03:20 usages according to the Torch Compile, so now they are faster. Moreover, with LoRA, now we support
00:03:27 FP8 Scaled, just as Musubi Tuner. This was also not a feature in the original Kohya SD Scripts;
00:03:35 now we are on-the-fly converting the base model into an FP8 Scaled model. And what
00:03:42 kind of impact could this be making? From a quality-wise perspective, this is making almost
00:03:48 no difference. So this is BF16 LoRA; this is FP8 Scaled weights LoRA. These are from grid images,
00:04:01 so these are not cherry-picked; this is BF16, FP8 Scaled. So there is almost no difference between
00:04:14 BF16 versus FP8 Scaled weights while training. What about the VRAM usage difference? So let's
00:04:24 load the configurations to see. Go to the LoRA tab, open all sections. Let's select our LoRA from
00:04:31 our LoRA configuration; the first config will be 24 GB Quality 1. You see this configuration? And
00:04:39 let's see the swap count, block swap count. So when we don't use FP8 Scaled, we have to do 13
00:04:48 block swap count. This means that it will use RAM for 13 blocks of the model. So it will be slow,
00:04:56 basically. And let's look at the 24 GB FP8 Scaled. When I load it and look at the block swap count,
00:05:04 it is zero. This means that you will get a huge amount of speed up if you have a 24 GB GPU.
00:05:12 Moreover, you see I do not have FP8 Scaled for 32 GB GPUs. Why? Because they are fitting into VRAM
00:05:22 completely right now since we have Torch Compile enabled in all our configurations, and it is even
00:05:29 reducing the VRAM usage further. These were LoRA configs. So in DreamBooth, sadly, you cannot train
00:05:36 models with FP8 Scaled; currently, only BF16 mixed precision training is supported. However,
00:05:43 don't worry, the 32 GB configuration is also fitting into the VRAM completely.
00:05:50 So for DreamBooth, for fine-tuning, don't forget to refresh, go to the DreamBooth tab,
00:05:55 and load from here; otherwise, it will not work. This is very important if you remember from the
00:06:01 original tutorial. The DreamBooth configurations go into here. This is also fine-tuning; do not use
00:06:06 this tab for fine-tuning. The DreamBooth tab is also for fine-tuning, and LoRA for the LoRA tab.
00:06:12 And what other new configurations do we have? When we look at the configurations, we now have an 80
00:06:18 GB GPU configuration. And what is this doing? This is speeding up the training significantly. If you
00:06:26 look at some of the training speeds which I have shared here—you see there are example speeds—you
00:06:33 can see that the 80 GB configuration on the RTX 6000 Pro is 1.7 seconds/it. This is for batch
00:06:41 size 1, 1024 by 1024 pixels. So if you train 28 images with 200 epochs, it is going to take 200
00:06:50 multiplied by 1.7, multiplied by 28. This many seconds, and this is going to take 158 minutes
00:06:59 total. And this will be the highest quality. And how much would this cost you on MassedCompute
00:07:05 with an RTX 6000 Pro? With our coupon, it would be 1.8 multiplied by 0.75, which is $1.35 per hour,
00:07:17 and multiply it by like 2.5 hours, it would cost you like $3.30 total. You can also use an RTX 4090
00:07:25 on RunPod; you see it is 2.66 seconds/it. So we have got amazing speed ups with our newest
00:07:33 configuration compared to our previous tutorial. Another new feature we have is in the Utilities
00:07:39 tab; you will see we have the FLUX FP8 Converter. This is for converting your FLUX DreamBooth models
00:07:46 into FP8, but this is also Scaled FP8, so the quality is amazing. What do you gain by this?
00:07:53 Normally, original models are 22.2 GB; with FP8 Scaled converted models, it is 11.1 GB. So
00:08:03 it will fit into your 12 GB GPUs as well. And do you lose quality? Let me show you. I have tested
00:08:10 all these configurations to verify them. I have grids for all of them. So this is here... okay,
00:08:17 let me open it, and let's uncheck all... and here: BF16, BF16 FP8 converted. So
00:08:24 the converted model versus BF16. So this is the original BF16 model, and this is the FP8 Scaled
00:08:32 converted model. This is the BF16 model; this is the FP8 Scaled model. You see, almost the same.
00:08:38 There is no quality loss when you convert your trained models into FP8 Scaled versions with our
00:08:55 newest tool. This is a very nice addition to our SECourses Premium Kohya GUI application.
00:09:03 Another amazing tool we have is Image Pre-processing. What this does is that
00:09:09 it pre-processes your training images exactly as Kohya is going to do during training. So you
00:09:16 will see your actual training images—how they were pre-processed, how they were processed. This will
00:09:22 help you to figure out your inaccurate images or orientation-problematic images. For example,
00:09:28 let's pre-process these images. Input images and output images will be "test pre," and I am going
00:09:34 to select FLUX; I am going to enable bucketing. You can also fix EXIF orientation to use them,
00:09:41 and these are our resolutions. Let's process images. So you see,
00:09:44 it is very fast; it is all pre-processed. When I go to this folder, you see:
00:09:50 this is how FLUX is going to use my images. When you open these images, you see it has added some
00:09:58 padding here; it has inaccurate orientation, so I would have very bad quality training if I had
00:10:06 used these images like this one. I mean, if I had used these images as my dataset,
00:10:13 they would be actually used like this, not like this. So this pre-processing tool will
00:10:19 give you a huge amount of information about your buckets because it will show you different aspect
00:10:24 ratio images as well in here, so I recommend you to use this for pre-processing your dataset to
00:10:31 see how they were actually used during training. What other options do we have as a new feature?
00:10:38 When you go to the presets, you will see that it is starting from 8 GB of GPUs. Moreover,
00:10:44 now I have optimized the model loading. So now we support memory-efficient loading;
00:10:51 therefore, you can do training with lower RAM memory as well—not only VRAM but also
00:10:57 RAM memory as well. Furthermore, when you start training right now, it will show you the actual
00:11:03 training speed after the very first step. This was problematic previously; now it is fixed. Moreover,
00:11:11 I have added CPU-based text encoder caching so that it will work with lower VRAM GPUs as well.
00:11:18 And we have a new model downloader. So when you download this zip file—let
00:11:23 me demonstrate for you—you will get to these files. It will contain the newest
00:11:28 configs, DreamBooth tab, LoRA tab; it will have test prompts that you can use, and it will have
00:11:35 "Windows Download Training Models." First of all, I need to start installation so it will generate
00:11:40 a virtual environment so that I can also use the Windows download models. When you run the Windows
00:11:46 download models, it will ask you which model you want to download. It will not download duplicates.
00:11:52 So you can download FLUX Dev FP8 Precision—I don't recommend this; this is not needed anymore.
00:11:59 You can download the regular FLUX Dev model as usual, as before. You can download the FLUX Krea
00:12:04 Dev model, but training on the Krea Dev model yields very low quality; I don't recommend it.
00:12:11 For realism, use FLUX SRPO Realism. If you want to use it for other things like stylization, 3D,
00:12:18 anime, use FLUX Dev. So you can directly select your option, and it will download
00:12:23 all the necessary models without any errors because this is also doing hash calculation,
00:12:28 hash verification, so they will be 100% accurate. And now let's make some comparisons with the FLUX
00:12:35 2 model because it was just published a few hours ago. Will it replace our existing FLUX
00:12:42 workflow? Currently, not yet, because the quality is not there yet. It is a very big
00:12:48 model. V1 is yielding better results, better quality with lower VRAM requirements, but I
00:12:55 will be following FLUX 2 very closely. So let's generate some images and let's make comparisons.
00:13:03 For example, this was generated with a FLUX SRPO trained model with myself. You see,
00:13:09 this is the quality; it is really good. This is 2048 pixels. I will copy the prompt. Then I
00:13:15 will generate it on base FLUX SRPO. For base FLUX SRPO, I am using the FLUX Dev models preset and
00:13:23 also using the upscale preset we have, this one. So let's copy-paste it and let's generate. Then
00:13:29 let's also generate 2048 by 2048 on the FLUX 2 model on the FLUX official website. So this is the
00:13:38 highest quality; this is the FLUX 2 Pro model, not the FLUX 2 Dev model, because we will have access
00:13:45 to the Dev model, not the Pro model. ComfyUI is already working on it,
00:13:48 and SwarmUI probably will add it after SwarmUI implements it. Hopefully, I will make a tutorial
00:13:55 and publish the presets as usual with automatic downloading with our downloader application,
00:14:01 so it will be so easy for you to use. Don't worry about that. Okay, this is FLUX SRPO;
00:14:06 it is being generated right now on my PC, and this is the FLUX 2 Pro model; it is also generating.
00:14:13 Okay, so this is the FLUX 2 2048 by 2048 pixels generation. I think it is pretty good quality in
00:14:24 my opinion, but this is the Pro model, not the Dev model, so we need to see it; we need to see
00:14:29 how much time it will take, how much memory it will use. And this is the FLUX SRPO model
00:14:36 for this prompt. Of course, FLUX 2 is expected to have better prompt following, but I can say
00:14:44 that this is a pretty good, pretty high-quality realistic image. And this is my fine-tuned model;
00:14:49 this is also excellent quality as you can see. This is myself, and this is the generated image.
00:14:56 Let's try another prompt and make a comparison. Okay, this one. So this
00:15:02 image has an unrealistic element. You see this animal? This is not a real animal; therefore,
00:15:08 you can see that it doesn't look as realistic as myself, as my image. Let's copy this prompt
00:15:14 and try with the SRPO base model. When you have things in your images, in your prompts,
00:15:19 that are not realistic, not existing in real life, the models will struggle to generate realistic
00:15:26 images. Let's see what FLUX 2 can do for this. By the way, if you use FLUX 2 on other services
00:15:33 like on Fal.ai or anywhere else, the quality is very bad compared to the FLUX 2 on BFL itself.
00:15:45 Moreover, set your resolution to the highest resolution; otherwise, the quality is low
00:15:50 again. So this is the maximum, very best quality of FLUX 2 models. Therefore, when we can locally
00:15:57 generate—hopefully maybe tomorrow, maybe today—we will be able to aim for the maximum quality. But
00:16:03 there is a problem because this would take a huge amount of time with 50 steps; currently,
00:16:08 this is probably what it is doing, maybe 20 steps. We need speed-up LoRAs to be able to use this
00:16:15 model fast; otherwise, it will take a lot of time. Okay, so this is the FLUX 2 model generation. I
00:16:24 think this is pretty cool; the image is realistic. You see this bird is not very realistic,
00:16:29 but this is expected because this doesn't exist in real life. So this is what we get,
00:16:34 but this is a really good image in my opinion. So this was my image. So this is FLUX 2 Pro;
00:16:41 this followed the prompt better perhaps, I am not sure. And this is my generated image.
00:16:48 So this was the prompt, you can read it, and this is the FLUX SRPO base model generation,
00:16:54 not fine-tuned. So my fine-tuned model has the most realism; FLUX SRPO is also very realistic,
00:17:00 and this is FLUX 2. FLUX 2 also has great realism. Let's test this prompt on Nano Banana as well. I
00:17:07 wonder what it can do, so I will make it image only. So this is Nano Banana Pro. Currently,
00:17:13 this doesn't generate very big resolution, but we will see. Okay, this is Nano Banana 2,
00:17:20 Nano Banana Pro generation. I think it is bad, not good. So you see it is not even very realistic,
00:17:27 so the FLUX 2 is better, but we cannot set higher quality with Nano Banana. I don't know, maybe it
00:17:32 can generate higher quality somewhere else, but this is its generation on a third-party provider.
00:17:39 I also want to test this on Seedream 4, so let's see what it can do. This is by default generating
00:17:47 very high resolution. This is a Chinese model as well, but not open source. So the FLUX 2 Dev model
00:17:54 with accurate settings is, we can say, promising. Hopefully, we will get this quality. If we can get
00:18:01 this quality, it is really good. I mean, this is a native generation probably; I don't know if they
00:18:06 are doing some upscaling or native generating. We are doing upscaling with this image with our
00:18:11 preset, but you can see that it has an excellent amount of details. And this is Seedream 4. You
00:18:17 see, Seedream 4 is also a really good model. I think, in my opinion, FLUX 2 is better.
00:18:23 Yeah, it is up to you, you can decide. You can also test on the BFL playground;
00:18:29 they give you 50 free images, so I am using them. And we already have our tutorials for, you know,
00:18:36 installing SwarmUI, using these things locally. So I hope you have enjoyed. Please like, subscribe,
00:18:44 and one more thing that I need to mention: we have got amazing quality with Q1 image models realism.
00:18:50 The tutorial is here. I think this is the king right now until we figure out how to fine-tune,
00:18:56 how to LoRA. So currently the king is Qwen image realism. Therefore,
00:19:02 I recommend you use Qwen image realism until we figure out how to do FLUX 2 training.
00:19:15 For example, this is from Qwen image realism; this is amazing. Watch that
00:19:19 tutorial to learn more about it. Hopefully see you later.
All reactions