Insights for the
work ahead.

AI, security, and technology delivery through a business lens.

Selected reading from Simon Willison, with summaries and a separate CumulusOps perspective on what it means for your team.

AI & ADOPTION

Source published

By Simon Willison

Prove an AI-built tool is ready for everyday work

Based on Vibe coding and agentic engineering are getting closer than I’d like

From the source

Reflecting on his own use of coding agents, Willison describes the temptation to review less code as the tools improve. He explores the growing importance of evidence from actual use and the confidence an organization needs before adopting software whose implementation is increasingly generated by agents.

CumulusOps perspective

A convincing demonstration is a starting point for an adoption decision. We recommend a bounded pilot with a named business owner, representative work, and an agreed decision date. Establish how the task performs today before introducing AI: time spent, errors, rework, and unresolved cases. That baseline gives the team a way to judge whether the new tool improves the whole process.

For example, an assistant that routes support requests could first suggest a destination while staff retain control. Compare its suggestions with actual resolutions, including ambiguous requests and tickets containing several problems. Count the effort needed to check and correct its work. Time saved on the initial classification has limited value if another team spends longer repairing the handoff.

Before expanding, set thresholds for acceptable error rates, review effort, operating cost, and recovery from failure. Keep a usable manual path and assign someone to handle exceptions. For a PE firm, a successful pilot at one company is evidence to examine before another rollout; differences in data, systems, and operating practices still need local validation.

AI & DELIVERY

Source published

By Simon Willison

Keep engineering discipline in AI-assisted delivery

Based on Writing about Agentic Engineering Patterns

From the source

Simon Willison introduces a growing collection of practices for software engineers using coding agents. He distinguishes deliberate engineering from unreviewed code generation, emphasizing professional judgment and automated testing as agents generate, run, and revise software.

CumulusOps perspective

For CumulusOps, useful AI delivery starts with an agreed business outcome and ends with a system someone can operate. A fast build still needs clear acceptance criteria, review of important decisions, and a support plan. We recommend evaluating a delivery partner on the evidence and handover they provide, alongside how quickly they produce a working demonstration.

An acquisition-related workflow illustrates the distinction. An agent might help compare application inventories from two businesses, but a useful result also identifies missing owners, conflicting records, and dependencies that require confirmation. Agree what it may infer, what it must flag as unknown, and which source records support each recommendation. People responsible for migration decisions should be able to trace and challenge the output.

Make the handover part of the original scope: where the code and configuration live, how credentials are managed, how changes are tested, and who owns failures after launch. Include a way to reverse a problematic release. A portfolio company should be able to maintain the implementation when the original developer or provider is no longer in the room.

AI & DELIVERY

Source published

By Simon Willison

Ask AI agents to demonstrate what they built

Based on Introducing Showboat and Rodney, so agents can demo what they’ve built

From the source

Willison introduces tools that help coding agents exercise software and produce a reviewable record of commands, results, and screenshots. His central point is that passing automated tests does not establish that a feature works in practice. Demonstrations give reviewers additional evidence and a way to spot missed behavior.

CumulusOps perspective

A delivery review should let the business see what the system actually does. We recommend a short, repeatable walkthrough using representative, non-sensitive examples, with the inputs, actions, and resulting records visible. Screenshots can support that record, but reviewers should also be able to inspect the resulting system state and repeat the important steps themselves.

Consider an agent that prepares a quote from a customer request. Demonstrate where product and pricing information comes from, how an incomplete request is handled, and where a person approves the proposed quote. Include an unavailable integration or rejected approval. Seeing the system stop safely and explain what needs attention can be as useful as seeing the successful path.

Keep the demonstration with the release it verifies, including the relevant configuration, limitations, and test results. Avoid placing customer data or credentials in the evidence. This gives operating teams a practical reference for training and handover, and a way to repeat the same checks after changes to the model, prompts, or connected business systems.

AI & TESTING

Source published

By Simon Willison

Test the business scenario, not just the generated code

Based on How StrongDM’s AI team build serious software without even looking at the code

From the source

Willison examines StrongDM’s highly automated development process, where agents work against specifications, independent scenarios, and simulations of external services. He focuses on how teams can verify software when agents write both code and tests. He also questions whether the approach’s substantial model costs make economic sense.

CumulusOps perspective

Business acceptance tests should come from the people accountable for the process. We recommend defining representative scenarios before implementation and keeping some separate from the examples used to develop the system. That creates a more useful challenge than asking an agent to prove its own assumptions with tests derived from the code it just wrote.

For an invoice intake workflow, scenarios could include a duplicate invoice, a supplier name that differs from the master record, an unreadable attachment, and a request to change payment details. Define the expected handling for each case, including when the workflow must stop for review. An output that looks plausible should not pass if it updates the wrong record or skips a required approval.

Start with the smallest test environment that exercises the important dependencies. A complete simulation of every connected platform may cost more to build and maintain than the initial use case warrants. Track successful outcomes, exception rates, processing time, and total cost, including human review. Expand testing around observed failures and repeat the critical scenarios whenever the implementation changes.

AI & SECURITY

Source published

By Simon Willison

Set clear boundaries before connecting AI agents

Based on The lethal trifecta for AI agents: private data, untrusted content, and external communication

From the source

Simon Willison describes the data exposure risk when an agent combines access to private information, exposure to untrusted content, and tools that can communicate externally. Malicious instructions in outside content can exploit that combination.

CumulusOps perspective

An AI integration deserves the same scrutiny as giving a new employee access to business systems. Our recommendation is to start with a map of the information it can read, the outside material it encounters, and the destinations it can send information to. In a PE portfolio, make those boundaries explicit for each company; common ownership should not automatically give an agent access across separate businesses.

Consider an assistant that prepares customer responses from support tickets. It may need the customer’s ticket history and approved product documentation. That does not establish a need for access to acquisition documents, every customer account, or unrestricted email delivery. Separate drafting from sending, restrict available destinations, and make the proposed recipient and content visible when approval is required.

Before launch, decide how to disable the integration, revoke its credentials, and investigate unexpected activity. Agree which events are monitored and who responds; adding AI-assisted monitoring does not remove the need for access controls. The useful design question is how much damage a compromised workflow could cause, and which enforceable boundary limits it.

AI & SECURITY

Source published

By Simon Willison

Give every agent a clear owner and limited permissions

Based on An Introduction to Google’s Approach to AI Agent Security

From the source

Willison examines Google’s approach to preventing unauthorized agent actions and sensitive-data disclosure. The framework emphasizes a human controller, limited permissions, confirmation for consequential steps, and auditable activity. His analysis questions whether models can reliably separate trusted instructions from outside content or safely decide their own permissions.

CumulusOps perspective

We recommend treating each production agent as an identifiable service with a business owner, a technical owner, and a documented purpose. Avoid giving it a shared administrator account simply because that makes integration easier. Define the accounts, records, and operations it needs, then enforce those limits in the connected systems and application controls.

For an employee onboarding assistant, drafting a checklist, creating an account, and granting privileged access are different levels of authority. Separate those operations and decide which require approval from an authorized person. The approval should identify the affected user, target system, and proposed access. The agent should not be able to approve its own request or broaden its own permissions.

Review access when an employee leaves, a process changes, or an acquisition introduces a new environment. Retain enough activity history to identify who requested an action, what was approved, and what actually happened, while limiting sensitive content in logs. Agree who can stop the service and revoke its access. These responsibilities should remain clear across CumulusOps, internal teams, and other providers.

AI & WORKFLOWS

Source published

By Simon Willison

Give AI the context to do useful work

Based on Here’s how I use LLMs to help me write code

From the source

Willison describes a hands-on development workflow: explore possible approaches, supply relevant context and examples, request a specific implementation, and then run and test it. His examples show an iterative process in which experience and judgment help identify mistakes and steer the model toward a useful result.

CumulusOps perspective

Custom AI scoping should uncover the knowledge that experienced employees use without writing it down. We recommend walking through a real task with the people doing the work: what starts it, which systems they consult, how they resolve conflicting information, and what makes the result acceptable. Capture examples of good output and cases that require judgment.

In an M&A integration, two application inventories may use different names for the same service or disagree about its owner. A useful AI brief would identify the approved sources, explain how conflicts should be presented, and require uncertain matches to remain visible. Giving the model more documents without deciding which records carry authority leaves an important business decision unresolved.

Turn that brief into a small implementation with limited access and a clear feedback loop. Have process owners review the output, refine the examples, and document remaining gaps. Assign ownership of the reference material as well as the software. If an operating procedure or system of record changes, the AI workflow needs a deliberate update and another check against representative work.

AI & ADOPTION

Source published

By Simon Willison

Build trust into the rollout of business AI

Based on Open challenges for AI engineering

From the source

In this annotated keynote, Willison examines obstacles to trustworthy AI adoption: unclear privacy communication, data exposure through prompt injection, misleading instructions in retrieved documents, and unreviewed generated content. He argues that people remain responsible for material they publish with AI assistance, even when the model produced it.

CumulusOps perspective

Trust in a business AI system depends on whether people can understand its limits and challenge its output. We recommend explaining the workflow in plain language before launch: what information it uses, which providers process it, how long information is retained, and who can access the results. Verify these details for the selected service and configuration rather than relying on a general privacy statement.

For an internal knowledge assistant, employees should be able to see the source behind an answer and recognize when information is missing or out of date. A confident response should not silently settle conflicting policies. Provide a route to the person who owns the process, and make reporting an incorrect answer part of normal use.

Introduce the system with examples of appropriate tasks and explicit boundaries for sensitive material. Assign review responsibility for customer-facing outputs and track corrections alongside adoption and time savings. Across a PE portfolio, shared training and evaluation practices can be useful, while access decisions and source material still need to reflect each company’s responsibilities and environment.

AI & SECURITY

Source published

By Simon Willison

Treat outside content as data, not instructions

Based on Prompt injection explained, November 2023 edition

From the source

In this 2023 explanation, Willison describes prompt injection as a weakness in applications built around language models. Instructions hidden in emails or web pages can redirect an assistant. When that assistant can access private information or operate tools, the consequences can include data disclosure or destructive actions.

CumulusOps perspective

We recommend including hostile content in the test plan for any AI workflow that reads material from outside its controlling team. That includes ordinary customer emails, supplier attachments, support tickets, and web pages. The business requirement is that this material can inform the task without gaining authority to change permissions, redirect information, or initiate unrelated actions.

Consider an assistant reviewing a vendor proposal. A document might contain text asking the assistant to disclose other bids or send internal notes to an external address. Test whether the workflow can attempt those actions, and whether controls outside the model block them. Use synthetic documents and test accounts so the exercise does not expose real commercial information.

Limit available tools and destinations, keep credentials out of model-visible content, and define what happens after suspicious behavior. Human review needs useful context: the requested action, destination, and information involved. Treat a passed attack test as evidence for that case, rather than proof of immunity. Repeat relevant tests after changes to document handling, integrations, or model behavior.

AI & ARCHITECTURE

Source published

By Simon Willison

Separate document reading from authority to act

Based on The Dual LLM pattern for building AI assistants that can resist prompt injection

From the source

Willison’s early proposal separates a model that follows trusted commands from a model that reads untrusted content without tools. He explains the risks of passing unsafe outputs between them and of social engineering. An April 2025 update points to CaMeL research addressing flaws in the original proposal.

CumulusOps perspective

The architectural lesson we would take into a business project is separation of responsibility. A component that reads outside material does not necessarily need permission to act on business systems. Start by asking which steps need AI judgment and which can follow explicit software rules. Keep the path from an interpretation to a consequential action narrow and reviewable.

For example, one stage could extract a supplier name, invoice number, and amount from an attachment into a defined record format. A separate stage could check that record against approved supplier data and purchasing rules before creating a draft for review. Free-form instructions extracted from the document should not become commands to the accounting system. Valid formatting alone does not establish that the contents are correct or authorized.

Adding a second model also adds complexity, cost, and another interface to test. We recommend using it only when the separation provides a clear benefit that simpler controls cannot provide. Document what crosses each boundary, how rejected or ambiguous records are handled, and which controls prevent the reading stage from acquiring execution privileges.

From the archive.

An earlier CumulusOps perspective, retained with its original date.

HIPAA Compliance and Cloud Platforms

One of CumulusOps's verticals is the secure management and migrating HIPAA PHI data into the cloud. HIPAA and HITECH require that covered entities and their business associates implement security measures to protect sensitive health data.

Archived content, originally published August 12, 2018. This historical note is not a statement of current compliance requirements or a certification of CumulusOps services.

Explore our current cybersecurity services

Turn an idea into a working plan.

Talk with us about your security priorities or a workflow you want to improve with AI.

Talk to an expert