auteur

Design and build cinematic, award-level web pages.

  • web-design
  • cinematic
  • scroll-animation
  • design-system
  • anti-slop
  • frontend

Declared platforms: linux · macos · windows

Install
npx skills add 'https://github.com/NousResearch/hermes-agent/tree/main/optional-skills/creative/auteur'
Download bundle ↓
main · 24fd22bScanned 2026-09-15

Contributors

GitHub-linked commit authors for this SKILL.md at the saved revision. Co-authors and history before file renames are not included.

File history ↗
View on GitHub

Auteur Skill

Ported from agiwhitelist/auteur (MIT), snapshot commit 9bca227d — see LICENSE. Scripts, templates and references are the upstream files (CRLF→LF), with Hermes adaptation notes and references/ path fixes as the only edits.

Auteur designs and builds web experiences the way a film director makes a film: script first, then assets, then the shoot, then the cut. It has three registers — build (an excellent conventional site), direct (a cinematic scroll-directed site) and system (a multi-screen product as one design system) — on one shared core of taste. Nothing ships until the page passes an executable anti-slop gate and the skill has looked at its own output.

When to Use

  • A landing page, marketing site, hero section, portfolio or product page has to be built or redesigned — and looking generic is not acceptable.
  • The brief asks for scroll animation, storytelling, or a site that feels like a film.
  • A product spans several screens that must feel like one thing — app, dashboard, admin, onboarding, docs.
  • Someone says make it beautiful, make it wow, cinematic, or design system, naming no technique.

Not for polishing a UI someone else built, and not for backend-only work.

Prerequisites

  • Node 18+ — every QA gate is a .mjs script run with node via the terminal tool.
  • Playwright (for the QA gates) — in the project directory: npm install playwright then npx playwright install chromium. Required by scripts/shoot.mjs, motionqa.mjs, systemscan.mjs, refscout.mjs, chromadiff.mjs, moodboard.mjs (the scripts import playwright at runtime; slopscan.mjs and source.mjs are dependency-light).
  • ffmpeg — optional; only for the video/score paths in references/assets.md and references/scroll-flight.md.
  • Hermes tools — use image_generate for image generation/editing, terminal for node/ffmpeg/npm, write_file/read_file for project files, vision_analyze to actually look at screenshots, and browser_exec for live-page inspection when a script isn't the right fit.

How to Run

What it actually does

  1. Commits the art direction in writing before any markup — one hue, one type system, a motion budget, named anti-references.
  2. Generates or sources the assets: Hermes' image_generate tool, Blender, depth maps, CC0 meshes and HDRIs with their licences recorded.
  3. Builds from proven recipes — one WebGL context, transform/opacity motion, scroll state machines.
  4. Gates the result: slopscan fails the build on concrete slop, motionqa fails it on dropped frames, systemscan fails it on cross-route drift.

Network access

The recon and sourcing scripts read live pages (awwwards, Bing/Pinterest/are.na image search, Poly Haven, Iconify, Google Fonts, Openverse, Coverr). Fetched content is treated as reference data and licence metadata — never executed, and no credentials, API keys or logins are involved. Skip phases 0–1 entirely if you don't want outbound requests; every other phase works offline. moodboard.mjs also downloads the image URLs returned by those search hosts to build the contact sheet.

Routing

Read the argument / brief and route:

  1. direct or the brief smells cinematic — "wow", "cinematic", "immersive", "storytelling", "launch page", "premium brand", "make people stop scrolling" → load references/direct.md and follow its phases. This is the flagship register.
  2. build or the brief is ONE conventional surface — a marketing page, a landing, a single product page → load references/build.md.
  3. system or the brief has more than one screen that must feel like one product — app, dashboard, admin, settings, onboarding, a docs or content site with real navigation → load references/system.md. The unit of design becomes the component × state, the failure mode becomes drift rather than boredom, and there is deliberately no peak. If you are already in build and a second screen appears, stop and switch: half a system is worse than either.
  4. edit or the request modifies a page this skill built (the project contains design/DESIGN.md) — "add a section", "change the pricing", "swap the hero copy" → read design/DESIGN.md FIRST and follow its Editing protocol: reuse its tokens, section-opening patterns, and motion families; after the change run slopscan and re-shoot the affected viewports. An edit that ignores DESIGN.md is a regression even if it looks good in isolation.
  5. recon <brief> or the ask is only for reference material — "find references", "put together a moodboard", "what's the state of the art for X sites" → load references/recon.md and run just that phase: scout live sites, build the moodboard, hand back design/refs/REFERENCES.md (with the steal: lines filled) and design/moodboard/contact-sheet.png (with the read filled). No commit-sheet, no build.
  6. audit <path-or-url> → load references/verify.md and run the verification pipeline on an auteur-built page. If the target is an existing UI auteur didn't build and the user wants it polished rather than rebuilt, say that a dedicated UI-polish/critique pass (upstream paired auteur with a separate 'impeccable' skill, not vendored here) is the right tool and offer to continue only if they want a rebuild.
  7. Ambiguous (e.g. plain "make a landing page") → ask exactly one question: "A great conventional landing page, or cinema mode with scroll direction and generated assets?" (upstream phrased these example briefs in Russian; translated here.) Then route. (Multi-screen briefs are not ambiguous — they are system.) Don't ask anything else yet — each register runs its own intake.

All three registers share phase zero, and its centre of gravity is the commit-sheet. Order differs: build runs recon → commit-sheet → mockup; direct runs recon → storyboard → commit-sheet → mockup, because the film's scenes are what the six decisions get made about; system runs recon → system-sheet (route map + component inventory) → commit-sheet → mockup, because the six decisions get made about a product, not a page. Either way nothing is coded before the sheet is full.

Working relationship with other skills

Auteur builds; it does not re-polish foreign UI. If the user has an existing interface that needs refinement, run a separate UI-critique pass (e.g. vision_analyze on screenshots plus the sibling design skills). Upstream paired auteur with an 'impeccable' critique skill (not vendored here); auteur's verify gate and an outside critique measure different things and coexist happily.

Quick Reference

The non-negotiables and the phase table. They apply to every register, every phase, always — even if no reference file has been loaded. Match-and-refuse: if you are about to produce one of these, stop and restructure the element.

Banned (rewrite, don't tweak)

#BanInstead
1border-left/border-right >1px as a colored accent on cards, callouts, alertsfull border, background tint, leading icon, or nothing
2Gradient text (background-clip: text + gradient)one solid color; emphasis via weight or size
3Glassmorphism as default (decorative backdrop-filter cards)rare and purposeful, or solid surfaces
4The hero-metric template (big number, small label, stat row, gradient accent)evidence in prose, one committed visual
5Identical card grids (same-size icon+heading+text, repeated)vary size, structure, or drop the cards entirely
6Eyebrow kickers (tiny uppercase tracked label) above every sectionone deliberate kicker max as a brand system; vary section openings
7Numbered section scaffolding (01 / 02 / 03) when order carries no meaningnumbers only for a real sequence
8Inter or Space Grotesk as the first font choicepick from a contrast-axis pair (see taste.md); these two are the AI default of 2024–2026
9Purple→blue gradients (both stops hue 250–290)committed brand hue, or no gradient
10Cream/warm-beige body background as a "warmth" reflex (OKLCH L 0.84–0.97, C <0.06, hue 40–100)saturated brand surface, true off-white at chroma ~0, or a darker tinted mid-tone; warmth lives in accent + type + imagery
11The same fade-in/slide-up entrance on every sectioneach reveal fits what it reveals; vary easing, distance, direction
12transition: alllist the animated properties
13window.addEventListener('scroll', ...)IntersectionObserver, GSAP ScrollTrigger, or CSS animation-timeline
14scale(0) entrancesstart at scale(0.95) + opacity
15Bento grids of near-identical or empty cells; white-card-on-white bentobento only with real visual variation per cell, else a different layout
16Copy tells: "Revolutionize", "Seamless", "Effortless", "Unleash", "Elevate", em-dash–heavy sentences, decoration strips like "BRAND. MOTION. SPATIAL."concrete claims in plain words
17More than one marquee per pageone, or none
18Instrument Serif / Playfair Display as the reflex "elegant serif"serifs chosen for the brand, not from the AI shortlist

A ban may be overridden only through a written auteur-allow (see Verification) with a real reason — a deliberate, argued choice is voice; a default is slop.

Critical numbers (memorize; full context in reference files)

  • Body text contrast ≥ 4.5:1 (large text ≥ 3:1). Placeholders too. Muted-gray-on-tinted-white is the #1 AI readability failure.
  • Body line length 65–75ch. Display heading ceiling: clamp max ≤ 6rem for headings in prose flow — a wordmark or a deliberately type-led hero is exempt and the commit-sheet must say so. Display letter-spacing ≥ −0.04em.
  • Durations: button 100–160ms · tooltip 125–200ms · dropdown 150–250ms · modal/drawer 200–500ms · any UI >300ms needs a written reason.
  • Enter/exit easing = ease-out. ease-in is banned on UI.
  • Animate only transform and opacity. Stagger 30–80ms.
  • Motion budget: ≤ 3 scroll-triggered pattern families per page; one primary wow peak, supporting scenes at lower intensity.
  • Scrub smoothing 0.3–0.8. Hero video ≤ 2MB. LCP < 2.5s. CLS < 0.1.
  • Fullscreen passes (bloom, grain, DoF, any full-frame shader) are priced per pixel, not per object — they, not geometry, are what blows the frame budget. A perf number counts only when measured at DPR 2 on a production build: DPR 1 quarters the cost of every such pass, and a dev server roughly doubles the frame.
  • prefers-reduced-motion = an alternative art direction (gentler, not zero), never an afterthought.
  • Content must be readable with JS disabled: reveals enhance an already-visible default, never gate visibility.

Phases at a glance

Phasebuild registerdirect registersystem registerReference to load
0recon → commit-sheet → hero mockup gaterecon → screenplay (STORYBOARD.md) → commit-sheet → hero mockup gaterecon → SYSTEM-SHEET.md (routes + component inventory + states) → commit-sheet → mockup gaterecon.md, then build.md / direct.md / system.md
1asset production (generate → edit → optimize)— (source icons/fonts via source.mjs)assets.md
2build the pageassemble the film (smooth scroll first, hero, scenes top-down)tokens → the shell → screens in traffic order → every statebuild.md / scroll-cinema.md / system.md + taste.md + motion.md
3verifyverify + CINEMA-QA.mdverify + systemscan across every routeverify.md
4lock the style: fill design/DESIGN.mdsamesame, but DESIGN.md is the component contracttemplates/DESIGN.md

The hero mockup gate (one static throwaway screen, screenshotted and approved before anything else is built) is the cheapest moment to change art direction — details in each register's reference. design/DESIGN.md is the style contract that makes every later edit stay in style (the edit route reads it first).

Never skip a gate because the intermediate result "looks done". The gates exist because a page that merely looks done is exactly what every other AI ships.

Procedure

The commit-sheet (before any code, both registers)

Slop is what happens when defaults make the decisions. The commit-sheet forces seven real decisions onto paper before the first line of code. Copy templates/COMMIT-SHEET.md into the project (e.g. design/COMMIT-SHEET.md) and fill all seven fields with non-defaults:

  1. Peak — the ONE primary wow moment (direct) or signature element (build). One sentence. If you can't name it, you're not ready to build.
  2. Color — primary as OKLCH + commitment tier (restrained / committed / full-palette / drenched) + one line: why this is not lavender, not cream, and not the category reflex + the background lightness as a number (target mean L), because "dark feels premium" is where this skill drifts, and a number can be checked afterwards where a mood cannot.
  3. Type — display + text pairing on a contrast axis (serif+sans, geometric+humanist, mono+serif...) + one line: why not Inter.
  4. Grid break — the one concrete thing that breaks the symmetric-grid default: an overlap, an asymmetric split, a diagonal flow, a full-bleed interruption. Name it specifically.
  5. Motion budget — how many scroll-pattern families (≤3) and what they are.
  6. Reflex check — write down: (a) what a generic AI would do for this category (first-order reflex), (b) what a generic AI avoiding (a) would do (second-order reflex — e.g. fintech → "terminal dark mode" is also saturated now), (c) your chosen deviation from both. If recon ran, (a) is not a guess: whatever design/refs/REFERENCES.md showed five times is the reflex, dated and with receipts.
  7. House tells broken — name the two (minimum) items from taste.md §2.5 you are deliberately not doing this time, and what replaces each. Fields 6a/6b are the reflexes of the category; these are the reflexes of this skill, which recur across unrelated projects and are invisible from inside any one of them: near-black backgrounds, mono service labels, the logo/status/action header, the scroll-instruction footer, amber-or-acid accents, the wordmark-as-hero, glow standing in for lighting. Measured across nine showcase builds, eight were dark and three landed within 0.002 of the same lightness. A tell that genuinely belongs here can stay — say why, as with an auteur-allow.

Gate: every field filled with a specific, non-default answer. An empty or generic field ("modern, clean look") means stop and decide. This artifact is checked again at verification.

Reference files

  • references/recon.mdphase 0 scouting, two executable legs: scripts/refscout.mjs profiles live award-level sites (real stack, pinned scenes, scroll budget, fonts, painted palette, screenshots — mechanics, not skins) and scripts/moodboard.mjs builds a numbered contact sheet from Bing / Pinterest / are.na so the art direction is decided from live material instead of memory. Also: query craft, the steal rule, how recon feeds the commit-sheet, and the "reference images are not assets" line. Load at the top of phase 0.
  • references/taste.md — the full anti-slop system: extended bans with replacements, second-order category reflex table, color strategy tiers, typography pairing, copy rules. Load for any visual decision-making.
  • references/motion.md — the motion school: when to animate, easing/duration/spring numbers, performance rules, motion budget, sound policy. Load before writing any animation.
  • references/build.md — the standard register process. Load when routed to build.
  • references/system.md — the multi-screen register: route map, the component inventory as a gate, the state matrix (empty/loading/error are not edge cases), density rules, the no-peak rule, and scripts/systemscan.mjs — which crawls every route, reads what the browser actually painted, fails a control type over its declared variant budget — counting states (disabled, current, inside a data-state row) separately, so implementing the state matrix never reads as drift — presses Tab to catch controls with no visible focus state, and renders one tile per rendered variant so drift is visible as well as counted. Load when routed to system.
  • references/direct.md — the cinematic register: screenplay contract, scene-sheets, dramaturgy, assembly order. Load when routed to direct.
  • references/assets.md — the media crew and routing (in Hermes: image_generate for all image generation and edits, terminal for ffmpeg/node; video via whatever image→video backend the user has), §0.5 source-vs-generate (scripts/source.mjs: CC0 glTF meshes, HDRIs and PBR materials from Poly Haven, icons, fonts, CC images, stock video — with a licence ledger, because generation cannot make geometry or an IBL and stock video must never be the peak), the consistency trick (edit frame A into frame B), local video via the first→last-frame chain, generated elements/mockups, the ambient score, the degradation ladder, and asset caching. Load during direct phase 1.
  • references/scroll-cinema.md — working code recipes: scroll-scrubbed video, canvas sequences, GSAP+Lenis foundation, CSS scroll-driven animations, text reveals, the two-keyframe WebGL displacement transition, view transitions, ambient audio, and the cinematic transition library (wipe, curtain, letterbox, shutter, depth parallax). Load during assembly.
  • references/scroll-flight.md — the video-scrub tier: a photoreal "fly through the world" hero driven by scroll, using the drop-in templates/scroll-flight-engine.js. The canonical recipe for scroll-scrubbed video (encode-for-scrubbing -g 8, encoded-frame posters, SSIM seam gate, chain architecture A/B, iOS/mobile decode hardening, crossfade-vs-seamless seams). Load when the hero should be photoreal footage/AI-video rather than real-time WebGL.
  • references/ambient-backgrounds.mdquiet texture for secondary sections and simpler builds (not a hero): a curated 6 editorial/analog effects (paper grain, ledger/blueprint rules, topographic contour, ink tide, sparse dust, one heat-haze shader) + a zero-motion static-mesh default. The governing rule (weaker than the quietest foreground element; one ambient per page), the CSS/SVG-first stack, and the feTurbulence-static perf rule. Load when a section needs to not be flat but must NOT compete with copy.
  • references/verify.md — the acceptance pipeline: slopscan → screenshot journey → motion/perf/audio QA (FPS at DPR 2 on a production build, long-tasks, audio-gate, reduced-motion, for Tier-1 scenes) → numeric rubric → reference diff (your frame beside the reference that set the direction, with scripts/chromadiff.mjs measuring the colour drift a model never sees in itself) → QA sign-off. Load at phase 3.

Weak-model note

If you are a smaller model executing this skill: follow the tables and numbers literally, fill every template field, run every gate command, and do not improvise beyond the reference recipes — the recipes are verified, your improvisation is not. When a reference file conflicts with your instinct, the reference file wins. Write files using paths relative to the project root; never retype an absolute path from memory (the skill's name "auteur" is one typo away from "author", and misspelled absolute paths scatter your output across the filesystem).

Pitfalls

  • Network recon: refscout.mjs, moodboard.mjs and source.mjs read live pages (awwwards, Bing/Pinterest/are.na image search, Poly Haven, Iconify, Google Fonts, Openverse, Coverr). Fetched content is reference data and licence metadata only — never execute it. Skip phases 0–1 to stay fully offline.
  • Different harness: these scripts and docs were written for a different agent harness (upstream drove asset generation through several local image CLIs). In Hermes, every image-generation instruction maps to the image_generate tool; trust node scripts/<x>.mjs --help output and actual node errors over doc prose if they drift.
  • Unverified commands: the scripts pass node --check syntax validation, but full runs (which need npm install playwright + a chromium download) were not executed during porting. Treat shoot.mjs, motionqa.mjs, systemscan.mjs, refscout.mjs, chromadiff.mjs, moodboard.mjs, source.mjs end-to-end behavior, and all ffmpeg/video-encode recipes, as unverified upstream claims until you run them yourself.
  • slopscan verified shape: node scripts/slopscan.mjs <dir> runs without npm deps; it prints per-rule findings and exits non-zero on failures (exit 0 when clean).
  • Font metadata cache: source.mjs font … caches Google Fonts' ~2.6MB metadata JSON as auteur-gf-metadata.json in the OS temp directory (os.tmpdir()), not the project; delete it there to force a refresh.

Verification

The page is not done when the code compiles. It is done when:

  1. node scripts/slopscan.mjs <src-dir> exits 0 (fails are fixed, not suppressed — /* auteur-allow: RULE_ID -- reason */ exists for deliberate choices and demands a real reason);
  2. node scripts/shoot.mjs <url> has produced screenshot journeys at 390 / 768 / 1440 and you have looked at every frame — text overflow, blank scenes, broken reveals, layout collapse are found by eyes, not by text search;
  3. the numeric rubric in references/verify.md passes (contrast, LCP, CLS, reduced-motion journey, scene variety);
  4. for direct register: CINEMA-QA.md (from templates) is filled with PASS on every row.

If any gate fails — fix and re-run. Report results honestly: "slopscan clean, 21 screenshots reviewed, LCP 1.9s" beats "looks great".

Discovery context

Discovered by repository scan. No exact path reference found in the snapshot’s root AGENTS.md.