back
98 comments
I've been using MiniMax H3 on my M5 Pro 64GB MacBook Pro through ComfyUI. It works extremely well.

I had to modify the default ComfyUI workflows to use a GGUF quant (city96's ComfyUI-GGUF custom node, UnetLoaderGGUF in place of the stock loader) [0].

I use the model labeled Q5_K_M. There is Q8_0 available as well, which is 34GB and fits fine in 64GB unified memory if you keep resolution modest.

The main issue is speed, a ~9-second 480x864 clip at 20 steps takes me a bit over an hour. So this will be cool to try for the speed up alone.

There's a lot of great information and workflows available to follow on the r/StableDiffusion subreddit.

[0] https://huggingface.co/Abiray/MiniMax-H3-GGUF/tree/main/unet

This implementation is much faster on my M5 Max, like a few minutes for the same video, but on an M5 Max with 128GB, didn't test on M5 Pro. About memory, could be executed on 64GB with a few changes.
GGUF is outdated in the latest versions of Comfy-UI. If you want a good balance of size, speed and quality you should use the int8_convrot model from the official Comfy Org Repo https://huggingface.co/Comfy-Org/MiniMax-H3/tree/main/diffus...
> a ~9-second 480x864 clip at 20 steps takes me a bit over an hour

that's rough. for comparison, i tried the exact same parameters on my 5090 RTX and it took 2 minutes to generate.

i believe diffusion models are primarily compute bound so the macs aren't really the ideal hardware for this kind of stuff

I wonder how much faster your m5 pro is compared to my M1 Max @ 64gb
What is the quality of the output like compared to something like Veo?
How much free space do you have left after running this llm model? Have you tried to develop your own model with the M5?
In the AMA Minimax said that H3 could support sparse attention, that would be a huge speedup! I wonder if there are any news on that. H3 is very cool. EDIT: testing a --sparse-attention optional mode based on what they said in the Reddit post.
This is where the DGX spark makes up a bit of the ground it loses on llm work, diffusion and cuda go together like peanut butter and jelly.
cough DiffusionGemma cough

Seriously, very dumb model compared to what you can run locally, but holy moly is it FAST on one GPU, seriously impressive. Can't wait for those to be scaled up a bit to fit perfectly within 96GB VRAM, then they'll be competitive.

On my 128GB M4 Max Mac Studio, generating a 15s 480p video with MiniMax H3 in ComfyUI takes an hour and a half.

Put Codex to work on deploying it now, hoping the speed can improve quite a lot :-) Thanks anyway

> On my 128GB M4 Max Mac Studio, generating a 15s 480p video with MiniMax H3 in ComfyUI takes an hour and a half.

That's crazy, a RTX Pro 6000 does that in in 2-3 minutes (give or take, depending on your exact settings). LLMs don't make the difference between standalone GPU vs unified memory + CPU so obvious as diffusion models seems to do.

Please keep up posted about the results!
> Put Codex to work on deploying it now

Which codex?

For gods sakes, Apple let people have run other GPUs instead of these pissweak 2012 class mobile GPUs
This still requires 128Gb of memory, right? Me and my lowly 96Gb, like a commoner; missing out on the fun.
From README:

> On the 128 GB M5 Max, clean end-to-end image+audio and embedded-video+audio renders completed in 74.58 and 76.99 seconds respectively, each with about a 40.1 GB peak physical footprint and zero swaps.

Looks like it uses 40GB? So your 96GB mac setup should work fine i guess (Model itself is 33B)

It shouldn't? Unless you're using BF16 for all weights (I'm using NVFP4 for the text encoder, otherwise everything BF16 (and audio F32)) you'll fit it all within 96GB VRAM, bugs non-with-standing :) I've been fitting this within 96GB VRAM without issues.
you should have a look at https://github.com/deepbeepmeep/Wan2GP which is the goto tool for "gpu poor", although as people below already pointed out you should be fine with comfyui's standard setup aswell
wow antirez does not sleep
when you have enough money to not have to worry about anything, you can go back to your hobbies. in this case, his hobby is programming.
Understatement of the year :-D
Wow, had to check some of his other repos; his the one behind dump1090
running this on an m1 max 64gb, it produces some really nice clips with music and audio, stitching the clips together after makes a nice short story generated entirely using a local ai system. It takes quite a while to generate each clip though, But at least it is working. I would like to know a few things:

-Optimal settings/configs examples for h3.c to help speed things up -Optimal recommended generation settings for each mac product, i am sure it is easy to do -Prompt generator assistant

Great work from the github author. The more i use it the more i realise i don't need a gui to generate video, just terminal.

My setup: Macbook #1 as a client Macbook #2 as server

-use macbook #1 terminal + ssh -run mactop in terminal tab to monitor Macbook #2's hardware during ai video generation -run h3.c on Macbook #2 via ssh session -transfer the output video file from macbook #2 to macbook #1 using terminal scp -view the video

Alright I’ve been afraid to ask but have been having trouble finding

What are some adult entertainment workflows in comfyui, I need best loras, best prompts to start with

and the communities, are they on telegram or something?

While you will find plenty of people willing to scam you to pay for “adult entertainment workflows”, the built in templates in ComfyUI for the model—perhaps dropping in a Lora Loader node for the a Turbo lora for speed—handle running the model, the subject matter adaptation isn’t really a workflow issue but one of reference/control images/audio/videos and prompting.

For Minimax H3, more than most models, you should read (and, if you are using an LLM for prompt assistance, make it sure it has access to) the official prompt guidelines, as each of the main models (fl2va that handles text-to-video and first- and/or last-frame-to-video and r2va that handles more complex reference cases) has its own structured prompt format (with many common features).

H3 is quite uncensored, but was not trained on p0rn, so it has no anatomy clues needed to generate that kind of stuff. For softer adult content it is reported to be fine on Reddit.
This is a healthy question. We want to use these tools for regular old human needs and desires.
There will Reddit subs for it though couldn’t tell you which off top of my head

I’d personally steer clear of messaging platforms for this - who knows what one might stumble into there

a friend told me there's a reddit called: unstable diffusion
How similar are Jeff Dean and Salvatore Sanfilippo?
My favorite Jeff Dean fact is that he’s also antirez. Which reminds me of my favorite Salvatore Sanfilippo fact. He’s also Jeff Dean
People are really good at stuff.

I noticed on a bar TV the other day that some of the Chromecast screensaver landscape photo credits were to Peter Norvig. They were really lovely pictures.

Anyone tried it with M4 Pro, 48GB of memory?
I'd love to know what the alternatives are and how this is better
This will be a little faster right now on an M4 or M5 because it's optimized for Apple Silicon. Assuming this model is still state of the art in six months, which might not be a surprise given how long other video models have taken, this should be much, much faster with the M7 chip.
I totally agree with you.
Noob question to all, is there any open source coding model that I can run on Mac mini 16gb?
I use Bonsai 27B ternary on my 24GB MBP. But, I believe you definitely can run it with 18 or even 8GB.
I'm looking to setup a way to create images for my own instagram marketing. I do not care how long it takes to make 10 variations of a post as that speed would still be faster than me making it.

Does this model work with ComfyUI easily? Can I just download it?

How long to generate a 10-sec 1920x1080 vid on Mac M4 64GiB?
I mean yeah, a 5090 is obviously gonna smoke a Mac lol. But that’s not really the point. If I can run this thing locally and get a video in a few minutes instead of an hour, that’s already pretty sick.
hey, this is nice!
neat — now I just need a machine with the memory bandwidth to render the three-second clip of my cat before the cat itself forgets what happened