Improve an Existing Agentic Pipeline

Problem Given an existing agentic LLM pipeline, identify its bottlenecks and weaknesses and propose concrete architectural improvements (HLD-style discussion).

Functional requirements

  • Decompose a user request into a plan and execute multi-step tool calls against it.
  • Carry context across steps without losing earlier intermediate results.
  • Recover from tool failures, malformed tool arguments, and model errors rather than aborting the run.
  • Validate intermediate outputs before they feed the next step.
  • Expose per-step traces so a failed run can be debugged after the fact.

Non-functional requirements

  • ~50-100 requests/sec at peak; a few million agent runs/day.
  • p95 end-to-end latency budget of a few seconds for interactive flows; async/queued execution for long multi-step runs.
  • Per-run token/cost ceiling, since every extra agent step multiplies LLM spend.
  • Graceful degradation when an upstream tool or model provider is down.

Key components

  • Planning/orchestration layer that decides the step sequence, instead of letting a single prompt improvise the whole run.
  • Memory and context management: summarize/compact history and retrieve only the relevant slice per step rather than replaying the full transcript.
  • Tool layer with strict schemas, retries with backoff, and per-tool fallback strategies.
  • Caching of deterministic intermediate results and repeated tool calls.
  • Guardrails/validation between steps — schema checks, sanity checks, refusal handling.
  • Observability: structured tracing of every step (prompt, tool call, output, latency, cost) plus offline evaluation against a fixed task suite.

Deep dives / trade-offs

  • Tool-calling reliability: strict schemas + validation + retries vs. letting the model self-correct; how many retries before falling back.
  • Latency vs. quality: how many planning/reflection steps justify the extra seconds and tokens; parallelizing independent tool calls.
  • Context management: long-context prompting vs. retrieval/summarization, and the accuracy lost when history is compacted.
asked …
LeaderboardSalaryAccount