Skip to main content

Ce site est aussi disponible en français.

Voir en français
AI agents

Building an AI Workflow: Start Small, Evaluate, Expand

Start with one task, representative inputs and a clear definition of an acceptable result. Then add the steps needed to produce it: retrieval, model calls, validation, tool use or human review. This article explains how to build and evaluate a focused workflow before introducing more complex orchestration.

By EpicflarePublished Updated 7 min read

Many operational tasks do not need an agent: an automated AI workflow is often enough. When an agent is justified, systems built from simple, composable patterns are easier to test and maintain. Start with the simplest version that could work, and add a multi-step workflow or an agent only when it measurably improves the result.

Define the Task Before the Architecture

Before writing prompts or choosing tools, write down what the workflow is for:

  • One task, described in a sentence, with its inputs and expected output.
  • A set of representative inputs, including difficult and unusual cases.
  • A definition of an acceptable result and how it will be checked.
  • What the workflow must not do, and when it should stop or hand over to a person.

Frameworks: A Choice to Evaluate

Frameworks such as LangGraph or Amazon Bedrock Agents can speed up a prototype. They also add layers of abstraction that can hide the underlying prompts and API calls, which makes debugging harder.

A framework can help with integration, state and monitoring, but also adds dependencies and conventions. Compare the work it removes with the constraints it introduces, using the requirements of the project.

Many of the patterns below can be implemented with a few direct calls to a model API. Whichever option you choose, make sure the team can see the prompts, model calls and tool calls the system actually makes.

Patterns, from Simple to Complex

Build your system by starting with the simplest pattern and adding complexity only as needed.

1. The Foundation: The Augmented LLM

This is the basic building block: a single model call enhanced with capabilities such as access to knowledge and playbooks, tools and memory. The model can generate its own search queries, select the right tool and decide what to remember.

  1. In
  2. LLM call

    • RetrievalQuery and results
    • ToolsCall and response
    • MemoryRead and write
  3. Out
The augmented LLM. A single model call that can query a retrieval source, call tools and read or write memory. More complex workflows are built by combining calls like this one.

2. Workflow: Prompt Chaining

A linear sequence of model calls where each step processes the output of the previous one. Think A→B→C. You can insert automated checks, or gates, between steps to maintain quality.

It suits tasks that can be broken down into fixed, sequential subtasks, like generating ad copy. Each additional step adds latency, so this is often the first trade-off to evaluate before considering more advanced behavior.

  1. In
  2. LLM call 1
  3. GatePass or exit
  4. LLM call 2
  5. LLM call 3
  6. Out
Prompt chaining. Each call processes the output of the previous one. A programmatic gate checks the intermediate result: if the check fails, the chain stops instead of passing the error on.

3. Workflow: Routing

An initial model call classifies an input and directs it to a specialized downstream prompt, tool or workflow. It is a common way to orchestrate several AI workflows or sub-agents.

When you have distinct categories of tasks that are better handled by specialized pipelines, routing avoids a one-size-fits-all prompt.

  1. In
  2. RouterClassifies the input
  3. One route selected

    • LLM call 1
    • LLM call 2
    • LLM call 3
  4. Out
Routing. A first call classifies the input and sends it to one specialized prompt, tool or workflow. Only the selected route runs.

Example

Routing automated campaign setup between performance campaigns, branding campaigns and social content posting.

Routing can also help control cost by sending simple queries to a smaller, faster model and complex ones to a more capable reasoning model. In each case, evaluate the selected model on representative inputs for that route.

4. Workflow: Parallelization

Some tasks in a workflow are heavy and time-consuming. Running several model calls at the same time and aggregating the results can reduce waiting time. Parallel calls are also useful when several outputs need to be produced or compared:

  • Sectioning: Breaking a task into independent subtasks and running them in parallel (e.g., writing different campaign reports and analyses that have no dependencies).
  • Voting: Running the same prompt several times to generate diverse outputs, then selecting the best one or looking for consensus (e.g., multiple code reviews to find vulnerabilities).
  1. In
  2. Run in parallel

    • LLM call 1
    • LLM call 2
    • LLM call 3
  3. AggregatorCombines or votes
  4. Out
Parallelization. Independent calls run at the same time. Their outputs are combined (sectioning) or compared to select a result (voting).

5. Workflow: Orchestrator-Workers

This pattern is related to routing. A manager model, the orchestrator, analyzes a complex task, breaks it into subtasks at run time and delegates them to workers: model calls, sub-agents or specific workflows. A final step synthesizes their results.

  1. In
  2. OrchestratorPlans subtasks
  3. Workers

    • LLM call 1
    • LLM call 2
    • LLM call 3
  4. SynthesizerMerges results
  5. Out
Orchestrator-workers. The orchestrator decides at run time which subtasks are needed and delegates them. Dashed links show that the number and nature of the subtasks are not fixed in advance.

6. Workflow: Evaluator-Optimizer

An iterative loop where one model call generates a response and another critiques it against a set of criteria (e.g., accuracy or goal achievement). The feedback is used to refine the response in the next cycle. It is useful when clear evaluation criteria exist, but difficult to implement reliably.

  1. In
  2. Repeat until accepted

    • GeneratorProposes a result
    • EvaluatorChecks the criteria
  3. OutAccepted result
Evaluator-optimizer. One call generates a response, another evaluates it against defined criteria and returns feedback. The loop ends when the result is accepted or a stop condition, such as a maximum number of iterations, is reached.

7. The Final Step: Autonomous Agents

If, after these approaches, the task still calls for an agent, you will need a model operating in a loop: it uses tools based on feedback from its environment to make progress toward a high-level goal. The agent plans, acts, observes the result and re-plans.

For open-ended problems where the solution path cannot be defined in advance, an agent can be appropriate, but it is also the most complex option to implement. It comes with higher costs and the risk of compounding errors. Always implement stopping conditions (e.g., a maximum number of iterations) so that the effort stays proportionate to the result.

  1. HumanGoal and review
  2. Action and feedback loop

    • LLM callChooses an action
    • EnvironmentTools and data
  3. StopGoal met or limit
Agent loop. The model chooses an action, the environment returns feedback, and the loop continues until the goal is met or a stop condition applies. A person can provide the goal, answer questions or review actions.

The ACI (Agent-Computer Interface)

Prompt engineering gets a lot of attention, and it is a critical skill to master, but it is not everything. How you define your tools matters as much as your main prompt. Think of this as designing an Agent-Computer Interface (ACI) with the same rigor you would apply to a Human-Computer Interface (HCI).

  • Put yourself in the model's shoes: Is it obvious how to use this tool from its name and description alone? Write tool descriptions as you would write a tutorial for a junior employee. Be explicit about format, edge cases and examples.
  • Make tools hard to misuse: Design the tool's arguments so that mistakes are harder to make, for example by enforcing formats and required context in your code, rather than leaving room for avoidable errors.
  • Test and iterate: Observe how the model uses your tools. Identify common mistakes and refine the tool's description or parameters to guide the model toward correct usage.

Evaluate, Then Expand

Run the workflow on your test set after each change and compare the results with the previous version. Track the share of acceptable results, the types of errors, and the cost and handling time per task. Add a step, a tool or an agent only when testing shows it is needed, and keep the simpler version as a baseline.

A Quick Reminder for Production-Ready AI Workflows and Agents

  • Embrace simplicity: Start with the simplest pattern that solves the problem. Don't add complexity unless you can measure the benefit.
  • Demand transparency: Make the plan and intermediate steps visible. This is crucial for debugging and for building user trust.
  • Define stopping conditions: Set limits on iterations, cost and actions, and define when the workflow hands over to a person.
  • Invest in your ACI: A clear, well-documented set of tools is the foundation of a reliable agent.

Go further

Agentic AI and media buying: An investment guide

A framework for assessing architecture, interoperability, operating costs and control when investing in AI for media buying.

PDF · 27 pages · Free · one short form unlocks every resource

Related reading