Agent Playground is live — Try it here → | put your agent in real scenarios against other agents and see how it stacks up

The Big Picture

Keep a supervisor's intent, the reference evidence, and a prioritized action list together so feedback is immediately actionable—leading to fewer clarifying questions and clearer tasks for junior artists.

Key Findings

Linking a supervisor’s goal to the exact reference region and to an executor-ready checklist makes reviews far more actionable. MOONWALK stores the brief, annotated references, the junior artist’s interpretation, the submitted work, and supervisor annotations in a single record so context carries across iterations. AI is used only to ask for missing context, surface the supporting evidence, and organize supervisor-authorized revision items—not to make aesthetic judgments. In studio testing with 19 professional practitioners, participants preferred MOONWALK over a chat-style interface and reported clearer, more traceable feedback. traceable record
Test your agentsValidate against real scenarios
Learn More

Key Data

174% of MOONWALK’s review outputs produced junior-executable checklists
263% reduction in senior–junior clarification needs (participants reported fewer follow-up questions)
3MOONWALK was rated above neutral on 12 of 13 measured usability and alignment items

What This Means

Tool builders and engineering leads designing collaboration features for creative studios will get a clear pattern to follow: preserve intent, evidence, and action in one traceable record. Studio supervisors and production managers can use the approach to reduce back-and-forth and hand off work that junior artists can act on immediately.

Key Figures

Figure 1 : Pre-production review in 2D animation and VFX often breaks the intent–evidence–action chain , as creative intent, supporting references, and revision rationale are lost across senior–junior handoffs in their Work in Process (WIP) artwork. MOONWALK maintains this chain by preserving intent, grounding review in inspectable evidence, and converting supervisor-authorized decisions into evidence-linked, executor-ready actions for revision items. The framework bounds AI to coordination—requesting missing evidence, reminding either role of stated requirements, and organizing authorized items—not to aesthetic judgment.
Fig 1: Figure 1 : Pre-production review in 2D animation and VFX often breaks the intent–evidence–action chain , as creative intent, supporting references, and revision rationale are lost across senior–junior handoffs in their Work in Process (WIP) artwork. MOONWALK maintains this chain by preserving intent, grounding review in inspectable evidence, and converting supervisor-authorized decisions into evidence-linked, executor-ready actions for revision items. The framework bounds AI to coordination—requesting missing evidence, reminding either role of stated requirements, and organizing authorized items—not to aesthetic judgment.
Figure 2 : Referencing in 2D pre-production. Artworks establish visual targets, while specifications and references bound the acceptable exemplar space. Current practice rarely preserves how a particular reference region or specification motivates a revision request.
Fig 2: Figure 2 : Referencing in 2D pre-production. Artworks establish visual targets, while specifications and references bound the acceptable exemplar space. Current practice rarely preserves how a particular reference region or specification motivates a revision request.
Figure 3 : MOONWALK interface panels across three review operations. Articulate Intent: (a) Spec/Brief records the active project goals, requirements, and constraints; (b) Reference Hub organizes visual exemplars with per-reference intent notes; and (c) Artist Interpretation captures the junior artist’s reading of the brief, references, and intentional deviations before review. Ground Evidence: (d) WIP Upload & Analysis evaluates the submitted work through eleven analysis dimensions against available project evidence; and (e) Compare & Gap Analysis presents candidate discrepancies together with their supporting specifications or references. Authorize Action: (f) Supervisor Review Canvas supports region-level annotation and review; and (g) Revision Consolidation organizes supervisor-authorized decisions into a prioritized, evidence-linked checklist for junior artists.
Fig 3: Figure 3 : MOONWALK interface panels across three review operations. Articulate Intent: (a) Spec/Brief records the active project goals, requirements, and constraints; (b) Reference Hub organizes visual exemplars with per-reference intent notes; and (c) Artist Interpretation captures the junior artist’s reading of the brief, references, and intentional deviations before review. Ground Evidence: (d) WIP Upload & Analysis evaluates the submitted work through eleven analysis dimensions against available project evidence; and (e) Compare & Gap Analysis presents candidate discrepancies together with their supporting specifications or references. Authorize Action: (f) Supervisor Review Canvas supports region-level annotation and review; and (g) Revision Consolidation organizes supervisor-authorized decisions into a prioritized, evidence-linked checklist for junior artists.
Figure 4 : MOONWALK system architecture across the three operations. Articulate Intent: the supervisor structures a Spec/Brief (DG1) and annotated Reference Hub (DG2). Ground Evidence: (a) eleven dimension agents score the WIP (eleven in total; six shown); non-visual factors, high-disagreement, low-score dimensions escalate to human review; remaining results feed (b) a three-model synthesis (GPT-4o-mini: visual observation; Gemini 2.0 Flash: specification and reference comparison; Claude 3.5 Sonnet: synthesis) yielding (c) a reference-grounded spec summary. Authorize Action: canvas annotations, client opinions, and analysis results merge in the (d) Revision Consolidation view, outputting a U1/U2/U3-prioritized, supervisor-authorized checklist (DG3).
Fig 4: Figure 4 : MOONWALK system architecture across the three operations. Articulate Intent: the supervisor structures a Spec/Brief (DG1) and annotated Reference Hub (DG2). Ground Evidence: (a) eleven dimension agents score the WIP (eleven in total; six shown); non-visual factors, high-disagreement, low-score dimensions escalate to human review; remaining results feed (b) a three-model synthesis (GPT-4o-mini: visual observation; Gemini 2.0 Flash: specification and reference comparison; Claude 3.5 Sonnet: synthesis) yielding (c) a reference-grounded spec summary. Authorize Action: canvas annotations, client opinions, and analysis results merge in the (d) Revision Consolidation view, outputting a U1/U2/U3-prioritized, supervisor-authorized checklist (DG3).

Ready to evaluate your AI agents?

Learn how ReputAgent helps teams build trustworthy AI through systematic evaluation.

Learn More

Yes, But...

Results come from formative (12 people) and summative (19 people) studies in two animation/VFX studio contexts, so findings may not generalize to all creative domains. The study does not isolate which individual interface or AI prompts drove benefits. Some senior users reported reduced perceived control and there is a risk teams may over-rely on the system for coordination rather than keeping interpretation conversations when needed. Evaluation-Driven Development (EDDOps) Inter-Agent Miscommunication

Deep Dive

MOONWALK reframes pre-production review around an explicit intent–evidence–action chain: record what the supervisor wants (intent), link that goal to the exact reference region or specification that supports it (evidence), and convert the decision into a prioritized, supervisor-approved checklist that a junior artist can execute (action). The interface keeps the brief, an annotated reference hub, the junior artist’s stated interpretation, the submitted work, supervisor annotations, and the revision checklist as one continuous project record. AI is intentionally limited to coordination tasks—asking for missing references, surfacing supporting evidence, comparing the submitted work to stated requirements, and organizing authorized revision items—while leaving aesthetic judgment to humans. intent–evidence–action chain The team ran a formative study with 12 practitioners to map common breakdowns, then a within-subject in-studio evaluation with 19 professionals comparing MOONWALK to a chat-only interface and to existing workflows. Participants favored MOONWALK on most measures: it produced 74% junior-executable checklists, reduced clarification needs by 63%, and was rated above neutral on 12 of 13 Likert items. Practitioners valued having specifications and references persist through review, being able to trace decisions back to evidence, and receiving executor-ready actions—but senior users sometimes felt the generated summaries were too diplomatic and noted limits on perceived control. The approach suggests a practical pattern for creative collaboration tools: preserve how intent is grounded in evidence and translate it into actionable tasks while keeping final aesthetic authority with humans. annotated reference hub
Not sure where to start?Get personalized recommendations
Learn More
Credibility Assessment:

Authors include recognized HCI/graphics researchers (e.g., Xiang Anthony Chen) suggesting established expertise. ArXiv venue and no affiliations listed keep it from a 5.