Skip to content

Discord Bot ​

mold includes a built-in Discord bot that connects to mold serve, allowing users to generate images and videos via slash commands.

Running ​

bash
# Server + bot in one process
MOLD_DISCORD_TOKEN="your-token" mold serve --discord

# Or run the bot separately (connects to a remote server)
MOLD_HOST=http://gpu-host:7680 MOLD_DISCORD_TOKEN="your-token" mold discord

Setup ​

  1. Create a Discord application at the Developer Portal
  2. Create a bot user and copy the token
  3. Invite with: https://discord.com/api/oauth2/authorize?client_id=YOUR_APP_ID&permissions=51200&scope=bot%20applications.commands (Send Messages, Attach Files, Embed Links + slash command registration)
  4. No privileged intents are needed (slash commands only)

Slash Commands ​

CommandDescription
/generateGenerate an image or video, including attachment-driven LTX-2 audio-to-video, retake, and keyframe modes and ordered MiniMax H3 references
/meshGenerate a Hunyuan3D GLB from one source image or semantic front/left/back/right multiview attachments
/transparentRender a subject on a transparent background (PNG or WebP with alpha), optionally cut out of up to three ordered reference images
/identityGenerate an image conditioned on a face reference photo (PuLID), with identity_strength and identity_start_step
/expandExpand a short prompt into detailed generation prompts
/remixRewrite a prompt into subject-preserving alternatives, one creative dimension each (dimensions, style, variations 1-5)
/upscaleUpscale an attached PNG, JPEG, or WebP with a server-side Real-ESRGAN model; the result is private to the requester
/searchSearch the external model catalog read-only; results are sanitized, bounded, and ephemeral
/modelsList available models with download/loaded status
/statusShow server health, queue summary, and every GPU/MIG device; large fleets paginate across limit-safe follow-up embeds
/quotaCheck your remaining daily generation quota
/queueInspect or control the host-global queue (guild-only, Manage Server, ephemeral)
/downloadsInspect or cancel host-global model downloads (guild-only, Manage Server, ephemeral)
/video-upscaleManage Library-backed video upscales (guild-only, Manage Server, ephemeral)
/galleryBrowse host-global Library metadata read-only (guild-only, Manage Server, ephemeral)
/admin reset-quotaReset a user's daily quota (requires Manage Server)
/admin blockTemporarily block a user from generating (requires Manage Server)
/admin unblockUnblock a previously blocked user (requires Manage Server)

Capability and exclusion matrix ​

Discord parity is defined by user capability, not by copying every CLI flag into a slash command. The bot is an HTTP-only client of one Mold server and uses that server's advertised generation profile as the authority for model availability, defaults, ranges, fixed controls, accepted inputs, and output formats; Discord does not define a separate availability allowlist. If an additive capability block is absent, the bot treats that as an older server and either uses the documented legacy behavior or gives a clear unsupported response.

The following matrix accounts for every current GenerateRequest field. "Server/default" means Discord deliberately leaves the field unset or sends a fixed transport value so the selected recipe or server remains authoritative.

GenerateRequest field(s)Discord statusDiscord surface or reason
prompt, modelImplemented/generate; focused commands also select a model and provide the prompt shape appropriate to that capability.
negative_promptImplemented/generate; none or - sends an explicit empty negative, while omission preserves the recipe default.
width, height, steps, guidance, seedImplemented/generate, /identity, and the applicable focused commands. Omitted values come from the server-advertised recipe.
batch_sizeIntentional exclusionDiscord creates one result per interaction. Batch authoring belongs to the CLI and Studio clients.
output_formatImplemented with delivery limits/generate chooses MP4 or GIF for video; /transparent chooses PNG or WebP; images default to PNG; mesh output is pinned to GLB. Other still containers remain CLI/API choices.
embed_metadataServer/defaultDiscord does not override the host's metadata policy.
offload, placementIntentional exclusionHost memory policy and component placement are operator and CLI/API controls.
scheduler, cfg_plusIntentional exclusionAdvanced sampler controls remain CLI/API controls.
source_image, source_image_nameImplemented/generate accepts a bounded PNG/JPEG attachment and records its sanitized upload name.
source_fitIntentional exclusionDiscord has no interactive crop/pad editor; it preserves source aspect through the shared fit helpers.
id_image, id_image_name, id_weight, id_start_stepImplemented/identity accepts one PNG/JPEG identity photo and exposes strength and start step.
id_images, id_image_names, true_cfg, cfg_start_stepIntentional exclusionMultiple identity photos and PuLID true-CFG controls remain CLI/API and Studio capabilities; /identity is already a focused, bounded Discord form.
edit_imagesImplemented with a Discord bound/generate exposes two ordered reference attachments; /transparent exposes three. The recipe may accept more through CLI/API.
reference_weightIntentional exclusionIP-Adapter reference strength remains a CLI/API and Studio control.
transparent_backgroundImplemented/transparent, offered only when the selected recipe advertises transparency and a compatible alpha format.
referencesImplemented for chat-suitable media/generate carries two ordered MiniMax H3 image/video/audio references; /mesh uses named image references for ordinary multiview generation. Mesh-reference workflows are excluded.
strengthImplemented/generate, with the selected recipe's source-strength semantics and admission bounds.
meshPartly implemented/mesh exposes ordinary image-to-mesh texture, octree_resolution, threshold, target_faces, and matting. texture_resolution, texture_view_count, and delight remain CLI/API and Studio controls.
mask_imageIntentional exclusionInpainting mask authoring needs a canvas/editor and remains in the CLI/API and Studio clients.
control_image, control_model, control_scaleIntentional exclusionControlNet setup remains a CLI/API and Studio capability.
expand, original_prompt, prompt_transformIntentional exclusion on generation requests/expand is a separate prompt tool. It returns text for review rather than silently transforming a generation request.
save_to_galleryServer/defaultDiscord generations follow the host's normal durable publication behavior.
title, tags, collectionIntentional exclusionLibrary filing and naming remain in the CLI/API and Studio clients.
batch_id, batch_index, batch_countIntentional exclusionThese are prepared-batch provenance fields, and Discord does not author batches.
mesh_workflowNot client-authorableThis is server-minted workflow provenance. Durable mesh workflows are CLI/API-only.
lora, lorasIntentional exclusionSingle and stacked LoRA authoring remain CLI/API and Studio controls.
frames, fpsImplemented/generate; duration is converted through the advertised temporal grid and limits.
upscale_modelIntentional exclusion on generation requestsChained post-generation upscale remains a CLI/API and Studio control. The separate /upscale command covers attachment-based image upscaling.
gif_previewImplemented transport policyThe bot requests a GIF preview for video so it can fall back when the primary MP4 exceeds Discord's upload ceiling.
enable_audioImplemented/generate exposes synchronized audio where the recipe supports it; omission preserves the recipe default.
video_onlyIntentional exclusionThe advanced LTX-2 video-only execution mode remains CLI/API.
audio_fileImplemented/generate attachment for LTX-2 audio-to-video.
audio_file_pathIntentional exclusionServer-local paths are trusted-deployment API/CLI inputs and are never accepted from chat.
source_video, retake_rangeImplemented/generate attachment plus start/end options for LTX-2 retake.
source_video_pathIntentional exclusionServer-local paths are trusted-deployment API/CLI inputs.
extend_video, extend_video_path, extend_overlap_framesIntentional exclusionVideo continuation and its server-local path form remain CLI/API and Studio capabilities.
keyframesImplemented with a Discord bound/generate accepts two or three LTX-2 keyframes, or the two endpoint images supported by Wan.
hdr_exr_dir, hdr_exr_full_floatIntentional exclusionHost-local EXR sidecar output is a CLI/API workflow.
pipelineImplemented/generate exposes ordinary LTX-2 recipes; attachment-driven audio, retake, and keyframe modes select their recipe automatically.
ic_lora_control, guidance_overridesIntentional exclusionLTX-2 camera/HDR adapters and detailed guider overrides remain CLI/API controls.
spatial_upscale, temporal_upscaleIntentional exclusionLTX-2 latent upscale controls remain CLI/API and Studio capabilities.
sample_shift, distill_strength_high, distill_strength_lowIntentional exclusionAdvanced Wan recipe controls remain CLI/API and Studio capabilities.

The profile-level controls follow the same boundary:

Generation profile contractDiscord handling
Recipe selector and defaultsReads the selected recipe's pipeline selector and defaults; focused attachments select specialized LTX-2 recipes.
Resolution domain, alignment, bounds, aspect, and presetsAccepts explicit width/height, derives a source/reference aspect where documented, and leaves final validation to the advertised recipe. Discord does not reproduce the Studio preset picker.
Step and guidance controlsExposes the numeric values and uses advertised defaults; fixed or hidden controls remain fixed or hidden.
Temporal frames, FPS, duration, and gridExposes frames/FPS plus a duration convenience option, snapping and bounding through the advertised temporal contract.
Prompt requirementHonors required, optional, or ignored prompt behavior. Ordinary /mesh needs image conditioning because Hunyuan3D ignores text.
Source image and strengthExposes both when the recipe accepts them and rejects incompatible combinations before upload where possible.
Negative prompt and audioExposes both when supported; omission preserves recipe policy.
Ordered reference imagesExposes ordered attachments within Discord's option and request-size limits; server admission owns the recipe ceiling and format rules.
Identity and transparencyExposed through /identity and /transparent, gated by server-advertised capability.
Mesh generationExposes ordinary image-to-mesh and named multiview controls through /mesh. Supplied-mesh, text-to-mesh, texture-only, mesh-roundtrip, and other durable workflow modes are CLI/API-only.
LoRA, ControlNet, masks, schedulers, sequence authoring, continuation, and advanced family controlsIntentionally excluded as listed above. Scripted chain jobs and their durable lifecycle are CLI/API-only.
Output formatsNarrows the advertised formats to containers Discord can deliver safely; the server still performs final admission.
Provenance and delivery notesDisplay metadata only; they are not user-editable controls.

Auto-chaining a single long video request is an engine implementation detail, not Discord sequence authoring. Discord has no /sequence, chain-job lifecycle, or durable mesh-workflow commands. Use mold run --script, mold jobs, mold mesh-workflow, or the corresponding HTTP APIs for those workflows.

Access boundary for host-global operations ​

Queue rows, download jobs, Library filenames, saved prompts, and video-upscale jobs belong to the Mold host; the current server contract does not attach a Discord user owner to them. /queue, /downloads, /video-upscale, and /gallery are therefore guild-only commands that require Discord's Manage Server permission and always reply ephemerally. Global queue pause/resume and per-item pause/resume are separate actions and call separate server endpoints; one never stands in for the other. Gallery access is read-only, and Discord does not expose destructive Library actions.

Operator commandSubcommands
/queuelist, show, global pause / resume, per-item pause-item / resume-item, cancel, and retry
/downloadslist and cancel; starting an install is deliberately absent
/video-upscalecreate, list, status, pause, resume, and cancel, using an exact Library filename as source
/galleryRead-only list and show; show may attach a bounded thumbnail

/search is a read-only user command. It applies the normal allowed-role check, bounds and sanitizes external catalog text, and replies ephemerally. Installing a catalog result or starting a model pull remains outside Discord so license acceptance stays explicit. An operator may cancel an already-created download with /downloads cancel; catalog license acceptance is never inferred from a search.

/upscale ​

/upscale accepts one PNG, JPEG, or WebP attachment up to 10 MiB, an optional upscaler model, and an optional tile size. It uses the same generation access, cooldown, daily-quota, and failure-refund policy as the rendering commands. The bot returns a PNG in an ephemeral response and refuses a result above its 24 MiB Discord delivery ceiling. This command calls the standalone image-upscale API; it does not set GenerateRequest.upscale_model on a new generation.

/identity ​

Face-identity conditioning has its own command rather than options on /generate. Discord caps a chat-input command at 25 options and /generate is already at exactly 25, and identity is qualified only for a fixed list of FLUX and SDXL checkpoints (see the Identity Photos guide), so none of /generate's video and conditioning options apply to it anyway.

OptionPurpose
promptRequired. What to render.
identityRequired. Face reference photo as a PNG or JPEG attachment.
modelIdentity-capable checkpoint. Autocompletes only models the server advertises; defaults to one.
identity_strength0.0–3.0, default 1.0.
identity_start_stepFirst identity-conditioned denoise step; must be under the resolved step count. Default 0.
width / heightOutput size in pixels.
stepsInference steps.
guidanceGuidance scale.
seedSeed for reproducibility.

Preconditions are checked in cost order, so an impossible request never takes a quota slot or a download: the declared attachment size and container and the strength range first, then the model gate against the server's advertised /api/models[].supports_identity (an absent field is read as "no", which covers both a server too old for identity conditioning and one whose binary cannot execute it) then the start step against the resolved step count, and finally the downloaded bytes. A server advertising no identity-capable model at all says so instead of guessing a checkpoint. The result embed carries an Identity row naming the photo, the strength, and the start step.

/transparent ​

Transparent-background renders have their own command for the same reason as /identity: /generate is already at Discord's 25-option ceiling. It offers only models whose recipe advertises capabilities.transparency (today Qwen Image 2.1; the model option autocompletes those, downloaded first) and sends transparent_background: true.

OptionPurpose
promptRequired. The subject alone — no scenery or backdrop.
modelTransparency-capable model; defaults to one the server advertises.
formatPNG (default) or WebP. JPEG is not offered because it cannot carry alpha.
reference_1 … reference_3Ordered reference images (PNG, JPEG or WebP), e.g. a picture to cut a subject from.
seedSeed for reproducibility.
width / heightOutput size. Without them the last reference's aspect ratio is used.
stepsInference steps.

The prompt you type is what is stored; the model's RGBA wording is added by the engine. reference_2 needs reference_1, and reference_3 needs both, so the order the prompt names ("image 1", "image 2") is never ambiguous.

/mesh ​

Mesh generation has its own compact command because /generate already uses Discord's 25-option ceiling. Attach source for a single-view Hunyuan3D model, or any non-empty subset of front, left, back, and right for a 2mv checkpoint. The two forms are mutually exclusive. If model is omitted, source selects the ordinary Hunyuan3D default and named views select the five-step hunyuan3d-2mv-turbo:fp16 tier. Optional seed, texture, octree, threshold, and target_faces values ride the same generation request and admission checks as every other client. The response includes the poster and GLB using the existing Discord mesh delivery policy.

Configuration ​

VariableDefaultDescription
MOLD_DISCORD_TOKEN--Bot token (required; falls back to DISCORD_TOKEN)
MOLD_HOSThttp://localhost:7680mold server URL
MOLD_DISCORD_COOLDOWN10Per-user cooldown (s)
MOLD_DISCORD_ALLOWED_ROLES--Comma-separated role names/IDs for access control (unset = all)
MOLD_DISCORD_DAILY_QUOTA--Max generations per user per UTC day (unset = unlimited; 0 = block all)

Video generation

Running /generate against a video model (ltx-video-*, ltx-2-*, wan*) produces an MP4 by default. Pass video_format: Animated GIF to receive a GIF instead. You can also attach a source_image for img2img on regular models, or as the first frame for LTX-2 image-to-video. When the rendered MP4 exceeds Discord's upload ceiling the bot falls back to the always-bundled GIF preview.

Use duration for the simple path: duration: 10 uses the selected model's default FPS and converts ten seconds to the nearest frame count on the selected model's advertised grid (8n+1 for the LTX families, 4n+1 for Wan). The bot uses the same advertised frame/FPS defaults as Studio, supports LTX-2's full 20-second single-generation limit, and rejects a duration beyond the selected model's limit. The existing frames and fps options remain available for precise control; duration and frames cannot be combined.

LTX-2 specialized modes are selected by their attachments: audio_file starts audio-to-video, source_video plus both retake times regenerates that time range, and two or three keyframe_* images are spaced across the requested frame count for interpolation. These modes are mutually exclusive in one command and do not need the pipeline option.

/generate's reference_1 and reference_2 attachments are routed by the selected model's capability, never its name. On a model that advertises capabilities.reference_images (FLUX.2, Qwen-Image-Edit, Qwen Image 2.1, and the SD 1.5 / SDXL image prompt) they become ordered edit_images; for Qwen Image 2.1 they may be PNG, JPEG or WebP, and without width/height the output takes the last reference's aspect ratio. Only an image-prompt model also keeps source_image; everywhere else a source image beside references is refused. For MiniMax H3 Ref2VA they are its ordered references, each an image, an H.264 MP4, or a WAV, and they cannot be combined with source_image, source_video, audio_file, the keyframe_* images, either retake time, or the pipeline option. Any other model refuses them with the server's own sentence. Ordering is explicit, so reference_1 must be present before reference_2.

Negative prompts: leaving negative_prompt unset applies the model's advertised default negative (Wan ships a tuned one). To explicitly disable the negative prompt, pass negative_prompt: none (case-insensitive; - also accepted); Discord's 25-option cap means the opt-out rides the existing option as a sentinel rather than its own toggle.

Re-register after command-option changes

/generate's prompt option changed from required to optional so an LTX-2 image-to-video run can be submitted with just a source_image. Discord caches command definitions, so the bot's slash commands must be re-registered after upgrading before users see the optional prompt or new duration option. The new /mesh command also appears only after re-registration. The prompt relaxation is guarded on both ends: the bot only skips the up-front check when visual conditioning (source_image, source_video, or keyframes) is attached, and the server's family-aware validator still rejects an empty prompt for every image family and for unconditioned text-to-video.

Block List

The /admin block command stores blocks in memory. Blocks clear when the bot restarts. For permanent restrictions, use role-based access via MOLD_DISCORD_ALLOWED_ROLES.

NixOS ​

nix
services.mold.discord = {
  enable = true;
  # tokenFile is loaded via systemd EnvironmentFile.
  # the file must contain: MOLD_DISCORD_TOKEN=your-token-here
  tokenFile = config.age.secrets.discord-token.path;
  moldHost = "http://localhost:7680";
  cooldownSeconds = 10;
};