DeepInfra raises $107M Series B to scale the inference cloud — read the announcement

inclusionAI/

Ming-Image-0.1-Design

$0.01

x (width / 1024) x (height / 1024) x (iters / 12)

Design-native, open-weight text-to-image (6B) — ranked #1 among open-weight models on Artificial Analysis's UI/UX Design leaderboard. Renders legible UI and poster text plus cohesive graphics, illustration and photography in ~1.7s. MIT-licensed.

Public
Zero retention
inclusionAI/Ming-Image-0.1-Design cover image
api

Input

Prompt

text prompt

You need to log in to use this model

Log In

Settings

Width

image width in px; must equal height (square canvas) (Default: 1024, 128 ≤ width ≤ 1920)

Height

image height in px; must equal width (square canvas) (Default: 1024, 128 ≤ height ≤ 1920)

Num Inference Steps

number of denoising steps (Default: 12, 1 ≤ num_inference_steps ≤ 50)

Guidance Scale

classifier-free guidance, higher means follow prompt more closely (Default: 1, 0 ≤ guidance_scale ≤ 20)

Seed

random seed, empty means random (Default: empty, 0 ≤ seed < 4294967296)

Output

generated image #0
Model Information

Ming-Image-0.1-Design

An open-weight, design-native text-to-image model (6B). It ranks #1 among open-weight models on Artificial Analysis's UI/UX Design leaderboard — it renders legible text (UI, posters, typography) and cohesive graphics, illustrations and photography, returning a 1:1 image in about 1.7 seconds.

Architecture

A Bailing mixture-of-experts multimodal LLM (prompt understanding) drives a diffusion transformer, decoded through a Qwen-Image VAE. BF16, single GPU. MIT-licensed (inclusionAI).

Inputs

FieldDefaultNotes
promptrequiredQuote exact strings to render them verbatim; say "absolutely no text" for pure visuals
width / height1024Square canvas (values should match)
num_inference_steps12Denoising steps
guidance_scale1.0Classifier-free guidance
seedrandomFor reproducibility

Output

One PNG image.

Prompting tips

This model rewards precise wording:

  • Put the exact strings you want in double quotes, and add "the only text is …".
  • Say "absolutely no text" for icons, illustrations and backgrounds.
  • Reuse identical style tokens across a set so the outputs look like one family.

Family & ecosystem

Part of the open-source Ming-Image 0.1 Design family: pair it with Ming-Image-0.1-Design-Layer to decompose a generated design into editable layers, plus the open Agent Skills Ling UI Design and Image-to-Editable-PPT for end-to-end design workflows.

License

MIT — open weights on Hugging Face and ModelScope.