<< All versions
Checklist for Feature Slices with
Skill v1.0.0
currentAutomated scan100/100saitarrun/devforge-ai/llm-evals
──Details
PublishedSeptember 28, 2026 at 04:39 AM
Content Hashsha256:9e804011bafaec66...
Git SHA
──Files
Files (1 file, 1.9 KB)
SKILL.md1.9 KBactive
SKILL.md · 37 lines · 1.9 KB
name: llm-evals description: Architecture standards, evaluation metrics, cost budget controls, and security guardrails for LLM, RAG, and Agentic features. version: 1.0.0
LLM & RAG Application Engineering Skill
This skill defines industry standards for developing, evaluating, and operating LLM, RAG, and AI Agent features cleanly.
Core Architectural Pillars
1. Evaluation & Benchmarking (Evals)
- Deterministic Evals: Assert output schema validity (Zod/Pydantic validation).
- Model-Graded Evals: Evaluate accuracy, grounding (faithfulness), hallucination rate, and context relevance using LLM-as-a-Judge.
- RAG Triad:
- Context Relevance (Retrieval quality)
- Groundedness (LLM stays within retrieved context)
- Answer Relevance (Output addresses user prompt directly)
2. Token & Cost Budget Guardrails
- Max Token Limits: Hard limit max response tokens per prompt.
- Circuit Breakers: Halt downstream requests if daily/hourly token expenditure exceeds set threshold.
- Semantic Caching: Store prompt/embedding responses in Redis/pgvector to eliminate redundant LLM calls.
3. AI Security & Safety
- Prompt Injection Defense: Sanitize user inputs; separate system instructions from untrusted user content.
- PII Scrubbing: Redact secrets, emails, SSNs, and credit card numbers before sending payloads to LLM APIs.
- OWASP Top 10 for LLMs: Guard against Insecure Output Handling, Excessive Agency, and Data Poisoning.
Checklist for Feature Slices with has_llm: true
- [ ] Schema validation for structured output (JSON mode / Tool Call response parsing)
- [ ] Fallback model strategy (e.g., fallback from primary model to secondary model on rate limit)
- [ ] OpenTelemetry LLM tracing (LangSmith / Helicone / Phoenix / OTel instrumentation)
- [ ] Token usage logging and latency tracking
- [ ] Eval dataset created with at least 10 gold-standard ground truth examples