// HACKER NEWS — CYBERSECURITY
H3-metal – Native MiniMax-H3 inference for Apple Silicon
Native MiniMax-H3 inference for Apple Silicon. The project is being built as a
sequence of working vertical slices: deterministic host/model metadata first,
then portable Metal block parity, prompt encoding, prompt-to-video/audio, and
first/last-frame conditioning and then ordered references.
Prompt-to-video/audio, first/last-frame conditioning, and ordered Ref2VA
image/video/audio references work end to end. The current work is incremental
H3-specific Metal performance and memory optimization on M3 Max and M5 Max.
The examples assume that the Hugging Face snapshot is in ./MiniMax-H3 and
that FFmpeg and FFprobe are available on PATH.
--info checks the model layout and prints the selected Metal device without
mapping all weights or generating media. Run ./h3 --help for the complete CLI
reference.
Without -p, the same binary starts an Iris-style interactive session:
Type a prompt to generate a numbered video. The session keeps the exact BF16
prompt conditioning, prepared DiT, and video decoder in memory, so repeating a
prompt with another seed avoids loading and encoding them again. Useful commands
are !status, !seed random, !seconds 2, !show, !save output.mp4, and
!cache. Use !help for the full, short list.
First/last-frame conditioning is persistent in the session:
Use !first clear or !last clear to remove an anchor. Generated videos are
written to the session directory printed at startup.
For a general Ref2VA conditioning image, use !ref-image PATH instead. Images
are appended in order and exposed to the model as , ,
and so on; filenames have no meaning to the model.
!refs lists the current order, !ref-remove N removes one entry, and
!refs clear removes them all. Ref2VA references cannot be mixed with
!first/!last anchors.