AI image requests get blocked by a layered moderation stack: a text classifier scans your prompt before generation, an image classifier scans the finished output, and category policies cover violence, sexual content, minors, real likenesses, and gore. These filters run deliberately broad, so benign prompts trigger false refusals — which you can usually route around by rephrasing.
Your request is being checked at two separate points, and either one can refuse it. Before any pixels are generated, a prompt classifier scans your text for flagged words and concepts — terms like "blood," "child," "weapon," or "dead body" can trip it even inside a legitimate context like a war-photography scene or a horror short. After generation, a second classifier scans the image itself, so even a clean prompt gets refused if the output accidentally contains gore-like texture, unintended nudity, or a face resembling a real public figure. That second layer is why the same prompt sometimes generates fine and sometimes gets blocked: the check is on the result, not just your words. OpenAI and Google both document this multi-stage design — prompt filtering, output filtering, and category policies — in their published safety guidance.
False positives are a normal failure mode, not a sign you did something wrong. Keyword matching is coarse: a period drama needing a funeral scene, a thriller needing an injury, or a documentary-style image of a teenager in distress all sit in categories the filters watch, even when the intent is narrative. Platform vendors acknowledge over-refusal as a known trade-off of running filters conservatively.
Thresholds also differ by model and platform — a refusal on one model is not a refusal everywhere. Working filmmakers document producing hard-hitting, non-PG cinematic content through invideo's model stack that other pipelines refuse outright; one creator put it as "invideo treats you like an adult." Inside invideo, the invideo agent routes image requests across Recraft, Nano Banana, and GPT-Image-2, so when one model declines a prompt you can re-run it on another without rebuilding your setup.
When a legitimate shot gets refused, work through these in order. First, rewrite the prompt in cinematographic language instead of naming the sensitive act: describe lighting, framing, shadow ratio, and atmosphere, and let implication carry the threat — dark-genre work is fully producible this way, as a documented ~90-second horror short shows with 400 video generations and 30 image generations of interrogation-room dread completed for $870. Second, strip or substitute the specific trigger word and re-submit; often one term, not the concept, caused the block. Third, for sensitive narrative imagery — documented cases include depicting deceased characters in vintage family photos — render the image in sepia tone and a low-detail style: the reduced explicitness passes the output classifier while preserving the story beat, and you can instruct the invideo agent to degrade the image further to match a period aesthetic. Fourth, switch image models, since each applies its own threshold. If all four fail, the content likely sits in a hard policy category (real people, minors in harmful contexts, explicit sexual content) that no rephrasing will clear — redesign the shot to imply rather than show.
Watch some of these to see what works for you:
invideo treats you like an adult and I can't stress this enough.
— a professional filmmaker comparing content moderation across AI video platforms