๐ทpig architecture gguf clip
- natively trained multipurpose clip for gk engine
- low vram alternative/substitute of t5xxl, umt5xxl, etc.
- 80-90% smaller, better efficient and cost effective (hopefully)
rationale
since LLM in diffusion model basically serves as text encoder, which means its primary goal/function is delivering a prompt to the generator precisely, no talk, no deep reasoning, no long response, make it as simple as possible
limitation/cutoff
we stop the training when similarity reaches .95 or over 20k cycles and similarity no less than .75
| adapter | stock | pig-clip | delta |
|---|---|---|---|
| t5 | 0.7872 | 0.7913 | +0.0040 |
| umt5 | 0.7533 | 0.7634 | +0.0100 |
| qwen3 | 0.9649 | 0.9671 | +0.0022 |
| qwen3vl | 0.9296 | 0.9378 | +0.0082 |
| qwen3vl (layerwise) | 0.9475 | 0.9504 | +0.0029 |
how it works
simply replace --t5xxl t5xxl.gguf with --llm pig_clip-nvfp4.gguf --llm-adapter pig_t5_adapter-f16.gguf
examples
test it with pixart
ggk diffuser engine -- --diffusion-model pixart-nvfp4.gguf --vae pig_pixart_vae_fp16-f16.gguf --llm pig_clip-nvfp4.gguf --llm-adapter pig_t5_adapter-f16.gguf -p "close-up portrait of dog" --diffusion-fa -v -o out.png
test it with sd-lite
ggk diffuser engine -- --diffusion-model model.gguf --vae vae.gguf --clip_l clip_l.gguf --clip_g clip_g.gguf --llm pig_clip-nvfp4.gguf --llm-adapter pig_t5_adapter-f16.gguf -p "close-up portrait of dog" --steps 8 --cfg-scale 1 --sampling-method euler --clip-on-cpu --diffusion-fa -v -o out.png
umt5 adapter
ggk diffuser engine -- -M vid_gen --diffusion-model wan2.1_t2v_1.3b-q4_0.gguf --vae pig_wan_vae_fp32-f16.gguf --llm pig_clip-nvfp4.gguf --llm-adapter pig_umt5_adapter-f16.gguf -p "a pig moving quickly in a beautiful winter scenery nature trees sunset tracking camera" --cfg-scale 6.0 --sampling-method euler -v -n "blurry ugly bad" -W 480 -H 480 --diffusion-fa --offload-to-cpu --video-frames 14 -o out.avi
qwen3vl-layerwise adapter
test it with krea2
ggk diffuser engine -- --diffusion-model krea2_turbo-q2_k.gguf --vae pig_wan_vae_fp32-f16.gguf --llm pig_clip-nvfp4.gguf --llm-adapter pig_qwen3vl_layerwise_adapter-f16.gguf -p "a lovely pig holding a sign says Hi" --cfg-scale 1.00 --steps 8 --offload-to-cpu --diffusion-fa -v -o out.png
qwen3-4b adapter
test it with x-image
ggk diffuser engine -- --diffusion-model x_image-nvfp4.gguf --vae pig_flux_vae_fp32-f16.gguf --llm pig_clip-nvfp4.gguf --llm-adapter pig_qwen3_4b_adapter-f16.gguf -p "cute anime style girl with pinky messy long hair blue eyes wearing a maid outfit with a long black gold leaf pattern dress and a white apron, it is a postcard held by a hand in front of a beautiful realistic city at sunset and there is cursive writing that says PIG" --cfg-scale 1.0 --steps 8 --offload-to-cpu --diffusion-fa -v -o out.png
qwen3vl-4b adapter
test it with mageflow
ggk diffuser engine -- --diffusion-model mageflow-edit-turbo-nvfp4.gguf --vae pig_mageflow_vae_fp32-f16.gguf --llm pig_clip-nvfp4.gguf --llm-adapter pig_qwen3vl_4b_adapter-f16.gguf --llm_vision mmproj-qwen3vl-4b-it-f16.gguf --ref-image sheep.png -p "a sheep in sunglasses" --cfg-scale 1.00 --steps 4 --sampling-method euler --diffusion-fa -v -o out.png
Reference
- Downloads last month
- 1,493
Hardware compatibility
Log In to add your hardware
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐ Ask for provider support
