Performance Optimization Skills
The evidence-driven loop that benchmarks a confirmed baseline and challenges it with candidates.
These skills form the optimization workflow: capture a workload contract, benchmark a confirmed
baseline, then challenge it with one candidate at a time until the Service Level Objectives (SLOs)
are met or the budget runs out. The loop also uses
deploy-dynamo-recipe,
listed on the Deployment and Operations page, to deploy the confirmed baseline and
each approved candidate.
A prompt that reaches them: “Optimize this deployment for output tokens per second per user under a 200 ms time-to-first-token SLO. Budget 8 GPU-hours and stop after three failed deployments.”
See the optimization loop for the full sequence and the evidence rules for benchmark validity requirements.