AI operations
David Robinson and OpenAI: Questions for AI Product Builders
A safety debate is a reason to ask sharper operational questions. This guide distinguishes the reported departure from the decisions teams can document in their own AI products.

David Robinson wrote in an October 3, 2026 Atlantic essay that he had resigned from OpenAI, where he led the writing of safety reports accompanying major releases. He criticized the company’s culture and the industry’s level of care. Those are his stated views; they should be attributed to him rather than presented as independently established findings.
For a team shipping an AI product, the practical follow-up is to examine its own decisions. What evidence supports deployment, which actions are automated, and who can stop the workflow when the result is wrong? The framework below is our application-level guidance, not Robinson’s checklist or an assessment of any company’s internal systems.
Start with the source and a clearly bounded claim
A personal statement can explain the author’s reasoning. It does not, by itself, measure the failure rate of a model in your application. Avoid turning a personnel story into either a blanket rejection of AI or an endorsement of a competing provider. Evaluate the task, the evidence, and the consequences of an error.
For internal discussion, create two columns: what the source actually says, and what your team proposes to do. Keeping those columns separate makes it easier to disagree constructively. A recommendation should have its own rationale, owner, and verification step.
Maintain a small deployment decision record
| Record | Example for a creative application |
|---|---|
| Purpose | Generate draft campaign visuals for a human editor |
| Allowed inputs | Approved product assets and the current campaign brief |
| Acceptance criteria | Correct product form, required copy, and suitable composition |
| Known limitations | Text or details may need correction after generation |
| Release owner | The editor who approves the final asset |
| Rollback | Disable automatic submission while preserving task history |
The record should be short enough that the team updates it when the model or workflow changes. Store representative successes and failures alongside it. A collection containing only attractive examples makes it hard to understand what the product will do with ordinary or difficult requests.
Revisit the decision when you add new inputs, broaden the audience, or allow the system to take a new action. Reusing the same model does not make those product changes equivalent. An assistant that drafts a message and one that sends it have different failure consequences.
Make SJolt API operations observable
For SJolt image and video tasks, your application receives a task identifier and queries for completion. Save that identifier with the model, request revision, and user-visible job before continuing. A finished task establishes that output is available; your application still needs to determine whether it meets the brief.
- Keep API credentials in the server environment and give the browser only the information needed to display progress.
- Separate a request timeout from a confirmed task failure. If you already have a task ID, check that task before creating another.
- Record the output selected for delivery and the person or rule that accepted it.
- Provide a clear failure state so a user can revise the request instead of waiting indefinitely.
Apply the same discipline to LLM outputs. Keep user instructions distinct from retrieved documents, and parse the actual response format rather than treating every returned block as display text. When a response is incomplete, make that state visible to the application.
Place review where it changes the outcome
Review is most useful at a concrete boundary: approving a brief, choosing a generated asset, or authorizing publication. Asking a person to approve every trivial intermediate step can create noise while the important decision remains hidden. Present the final candidate with the evidence needed to judge it.
For example, show a proposed product image next to its approved reference and a short list of requested edits. An editor can then check whether the package shape and label survived. A generic success badge cannot answer those questions. If the reviewer rejects the result, preserve the reason as a specific instruction for the next attempt.
Choose a small number of operational signals that can trigger investigation: repeated task failures, unusual submission volume, missing output, or a rise in rejected assets. Implement those checks in your own application according to its needs; this article does not claim they are automatic SJolt features.
Run one failure exercise with the team
Take a representative job and interrupt it after submission. Ask another person to recover its state using only the saved record. Then supply a result that is technically complete but visibly wrong. Check whether the system and the reviewer distinguish those two cases.
The exercise should end with an observable improvement: a saved identifier, a clearer error message, a review screen, or a documented stop procedure. These are concrete ways to make an AI application easier to understand and correct, regardless of the model provider you choose.
Sources & further reading
Take the next idea into production.
Explore the models, test a workflow in the playground, and use the same request in your application.