Agent workflows
OpenAI Model Plays World of Warcraft: How the Agent Worked
A Warcraft demonstration is most useful when you understand what the model could observe and control. Here is the experiment behind the headline and a practical workflow design exercise.

The OpenAI model playing World of Warcraft in this demonstration was GPT-6 Astra, running through Codex and the independent agent-wow project. The project author reports completing the orc starting-zone quests in 40 minutes with no deaths. That is an account of one experiment, rather than an independently reproduced performance benchmark.
The interesting engineering question is how a language model can observe a changing world, choose an action, and determine whether that action worked. That same loop matters when an application uses an LLM to coordinate image or video tasks on SJolt.
What the demonstration actually used
In an October 2, 2026 write-up, the author describes using GPT-6 Astra at xhigh reasoning effort. agent-wow interacts with the game through its network protocol instead of feeding rendered frames to a vision model. The experiment ran on a private local AzerothCore server, not the live World of Warcraft service. The author also describes extracting quest information from server source files.
Those conditions define the result. A benchmark based on screenshots, a restricted action menu, or an unfamiliar map would measure a different system. The model, available information, tools, and environment must be considered together. A successful starting-zone run also does not establish performance on every quest, raid, or multiplayer situation.
Separate the model from the application around it
| Layer | Question to ask | Media-workflow equivalent |
|---|---|---|
| Observation | What information is visible at this step? | The approved brief, uploaded references, and current task state |
| Decision | What should happen next? | Choose the next shot or request a specific revision |
| Action | Which operations can the application execute? | Submit a supported image or video request |
| Verification | What evidence confirms completion? | A successful task response plus inspection of the output |
Treat these as separate interfaces. A model may produce a plausible plan while the action fails because an asset is unavailable. Conversely, an API may successfully produce a video that misses the creative brief. Keeping operational success and creative acceptance separate makes both failures easier to diagnose.
Give an agent a compact, explicit state
For an original game trailer, a useful state might contain the shot identifier, approved character references, the last submitted task ID, and the acceptance decision. The agent should not have to reconstruct those facts from a long conversation every time it resumes. Keep the state in your application and supply the relevant portion for each decision.
shot: courtyard-arrival
brief_status: approved
reference_asset: hero-v3
generation_task: saved-task-id
render_status: running
next_action: query_existing_task
acceptance: pending_visual_reviewThis is an application design example, not an SJolt API payload. Its purpose is to show why a saved identifier matters: a worker that loses its place should first recover the existing operation, rather than create another generation. Store human corrections alongside the brief so that retries do not revert to an obsolete instruction.
Apply the loop to a SJolt media pipeline
SJolt exposes GPT-6 Astra for text and image analysis, and separate media models for generation. Your application can ask the LLM to turn an approved brief into shot descriptions, then validate those descriptions against the selected media model before submitting them. Calling the Astra endpoint alone does not provide the agent-wow environment or operate a game.
- Start with one original shot and a concrete acceptance rule, such as a stable character appearance and an unobstructed closing frame.
- Choose the media endpoint for the actual job: a new scene, a first-frame animation, or an edit of existing footage.
- Save the returned generation task ID, query its state, and retrieve output URLs only after success.
- Review the complete output. Accept it, request one specific change, or stop the workflow with a clear reason.
Measure repeatability before calling it autonomous
Build a small evaluation with several briefs of different difficulty. Record completed objectives, interventions, failed requests, accepted outputs, elapsed time, and total spend. Report the denominator: five successful outputs from five attempts means something different from five selected outputs out of fifty.
Also include a deliberately interrupted run. Restart the worker after submission but before completion and check whether it recovers the saved task. Give it an unavailable reference and check whether it reports the missing input. These tests reveal workflow reliability more directly than another impressive final clip.
The Warcraft example is a useful prompt for designing better interfaces. For a SJolt application, the practical goal is a workflow whose next step and evidence are understandable to the person reviewing it. Build that observability into the first prototype, while the number of actions is still small.
Sources & further reading
Take the next idea into production.
Explore the models, test a workflow in the playground, and use the same request in your application.