Free agent skill · MIT

GPT Image 2 prompt skill for Claude Code, Codex and Cursor

Drop one folder into your agent’s skills directory and it gains 22 industrial GPT Image 2 prompt templates across 13 categories — with the style tags, scene tags, guidance and failure modes that decide whether a generated image is usable. No API key, no network calls, nothing to run.

What the skill does

Most image prompts fail for structural reasons, not creative ones: the aspect ratio is never stated, the on-image text is left to the model to invent, the layout hierarchy is implied rather than described, and the things that must not appear are never named. The skill fixes that by making your agent pick a template first and fill it second.

When you ask for an image prompt, the agent reads the bundled template library, matches your request to a category, then to a visual style tag, then to a scene tag, and either commits to the strongest template or offers you two or three with one-line reasons. It fills every placeholder with your specifics and returns a prompt built from six blocks: subject and task, composition and layout, visual style and materials, text and label requirements, aspect ratio and output format, and constraints and negative details. It closes by naming the template it used and linking the page where you can see how other people applied it.

It writes prompts; it does not generate images. Paste the result into whichever image model you use. The templates are tuned for GPT Image 2 and GPT Image 2.5, where accurate text rendering and layout control make the constraint blocks pay off, but the same structure transfers to other current models.

What is in the download

  • SKILL.md — the trigger description and the selection workflow the agent follows.
  • references/template-library.md — all 22 templates in full: category, styles, scenes, tags, when to use, guidance, pitfalls, a copy-ready text block, an agent-friendly JSON block, and links to real example cases in the gallery.
  • agents/openai.yaml — display name and default prompt for Codex-style hosts.
  • LICENSE.md — MIT terms and upstream attribution.

The whole package is plain Markdown and YAML — roughly 50 KB unzipped, with no scripts and no dependencies. You can read every byte of it before you install it.

Install the skill (available now)

Claude Code

~/.claude/skills/ (all projects) or .claude/skills/ (one project)

  1. Download the zip and unzip it, keeping the picgens-gpt-image-2-prompts/ folder intact.
  2. Move the folder into ~/.claude/skills/ to make it available everywhere, or into .claude/skills/ inside a repo to scope it to that project.
  3. Start a new Claude Code session. The skill loads when you ask for a GPT Image 2 prompt, an image-generation template, or a visual style.

Codex

~/.codex/skills/

  1. Unzip the download and move picgens-gpt-image-2-prompts/ into ~/.codex/skills/.
  2. agents/openai.yaml ships in the folder, so the skill shows its display name and default prompt without extra setup.
  3. Reopen Codex and ask it to write a GPT Image 2 prompt.

Cursor, Windsurf, and other editors

Project rules

  1. Open SKILL.md and paste its body into a project rule (Cursor: .cursor/rules/, Windsurf: .windsurfrules).
  2. Add references/template-library.md to the same rules folder, or attach it to the chat when you need the full template text.
  3. Reference the template library by name in your prompt so the assistant reads it before answering.

Any other agent

Context file

  1. Serve or attach SKILL.md as system context — it is plain Markdown with YAML frontmatter and no runtime dependencies.
  2. Attach references/template-library.md when the agent needs the full template text rather than the workflow.

A worked example

Say you ask your agent: "Make me a diagram that explains how a city actually works underneath the street — for a talk, portrait orientation."

There is no style word in that sentence and no aspect ratio, but there is a deliverable — an explanatory diagram — so the skill lands in the Charts & Infographics category and picks the Infographic Engine template. That template’s guidance is specific: define 3-5 modules, information flow, visual hierarchy, and short labels. Its recorded pitfall is equally specific: avoid long paragraphs inside the image. “For a talk” becomes a legibility constraint; “portrait” becomes 9:16. The agent hands back something like this:

Vertical 9:16 isometric cutaway infographic titled "Urban Metabolism Atlas".
Subject: one city block sliced from sky to bedrock — towers, street level, subway tunnel, utility trench, groundwater.
Modules (5, top to bottom): Power, Water, Transit, Waste, Data. Each module gets a short label, one icon and one flow arrow.
Composition: single continuous cutaway, left-to-right reading order inside each band, generous margins.
Style: clean technical illustration, muted slate base with one accent colour per module, thin consistent line weight.
Text: only the five module labels and the title, in English, large enough to read on a projected slide. No paragraphs, no lorem ipsum, no garbled glyphs.
Output: 9:16, high detail, flat lighting, no photographic depth of field.

Note what the model is no longer left to guess: the module count, the reading order, the exact strings that may appear as text, the line weight, and the two artifacts to avoid. That is the whole value of a template — it remembers the constraints you would only have discovered after three bad generations.

How it relates to the gallery

The skill is the offline half of the PicGens template library. The same 22 templates are browsable on the web at /templates, where each one shows example images, every prompt variant, and the cases it was derived from — useful when you want to look before you commit, or share a reference with someone who is not running an agent.

MCP server (in development)

Where the skill ships a fixed library as files, the PicGens MCP server will let an agent query the live gallery at request time — so a prompt written this week is available to your agent this week. The two halves are complementary: the skill decides how to write a prompt, the server supplies which prompts exist right now. The skill needs no key and keeps working either way.

These are the tools it is being built around:

  • search_prompts(query, model?, tag?, limit?)

    Full-text search across the live gallery, returning prompt text, model, tags, creator attribution and the source link for each hit.

  • get_prompt(id)

    Fetch one prompt by its PicGens id, with the complete prompt text, image metadata and the canonical page URL to cite.

  • list_templates(category?, tag?)

    List the 22 prompt templates with category, tags and a one-line description, so an agent can choose before it fetches.

  • get_template(slug)

    Fetch one template in full: when to use it, guidance, pitfalls, the copy-ready text block, the JSON block and linked example cases.

Status: in development — API keys issued on the account page will work with it at launch.

If you are signed in, you can create a key now on the API Keys page and keep it for launch. Nothing on PicGens requires one today: the gallery, the template pages and everything in the next section are open.

Machine-readable resources available today

You do not have to wait for the server to point an agent at PicGens. Everything below is public, unauthenticated and stable:

GET /api/templates returns an object with count, an attribution string and a templates array. Each entry carries the slug, its page URL, the English title and description, category, tags, styles and scenes, the “use when” line, guidance and pitfalls as arrays, the copy-ready text block, the JSON block where the template has one, and the example cases with links back into the gallery:

curl -s https://www.picgens.com/api/templates | jq '.templates[0].slug'

Frequently asked questions

What is an agent skill?
A skill is a folder containing a SKILL.md file with a name and a description, plus optional reference files. Agents such as Claude Code and Codex scan their skills folder, read each description, and load the full instructions only when a request matches. This one loads when you ask for a GPT Image 2 prompt, an image template, or a visual style.
Does the skill generate images?
No. It writes prompts. The skill turns a rough intent into a structured GPT Image 2 prompt built from 22 templates, then hands you the text to paste into whichever image model you use. It needs no API key and makes no network calls.
Which models does it work with?
The templates are written for GPT Image 2 and GPT Image 2.5, where they take advantage of the model’s text rendering and layout control. The same structure transfers to other current image models, though on-image text fidelity varies.
Is it free, and can I use the prompts commercially?
The skill is free and MIT-licensed. The template text derives from the MIT-licensed awesome-gpt-image-2 catalog by freestylefly. Prompts you write with it are yours. Example images stay with their original creators — a link from an example case is not a licence to reuse that image.
How does this relate to the PicGens MCP server?
The skill is offline and static: it ships the template library as files. The MCP server, which is in development, will let an agent search the live PicGens gallery and pull fresh prompts at request time. Keys created on the API Keys page are for that server and will work with it at launch; the skill needs no key and works today.
How do I update it?
Download the zip again and replace the folder. The skill carries no version state, so replacing the files is the whole upgrade. The template library is regenerated whenever the upstream catalog changes.

Keep exploring

Template copy and prompt text derive from the MIT-licensed catalog awesome-gpt-image-2 by freestylefly. Packaging and links by PicGens; see the bundled licence. Example images stay with their original creators — a link is not a licence to reuse an image.