mirror of
https://github.com/magnus919/agent-skills.git
synced 2026-09-17 06:26:31 +03:00
Greenfield SkillOpt: 3 epochs optimizing discoverability, decision guidance, and pattern expansion for a brand-new LlamaIndex skill. Epoch 1 — Prominence & Navigation: - Moved Key Principles to top (before Quick Reference) - Added Phase column to Quick Reference in workflow order - Added 'When to use' column to Template Files - Replaced flat When NOT to Use with Framework Routing Guide (LlamaIndex vs LangGraph vs PydanticAI vs Haystack vs DSPy) Epoch 2 — Decision Guidance: - Added Where to Start table (maps existing work → pipeline entry point) - Added Pipeline Mode table (Quick/Full/Evaluate/Graph modes) Epoch 3 — Pattern Expansion: - Replaced flat gotchas with structured Troubleshooting Recovery Guide (3 symptom categories with immediate + permanent fixes) - Added references/example-rag-pipeline.md (end-to-end worked example) - Added references/evaluation-workflow.md (ParamTuner + evaluators) 17 files: SKILL.md, 11 references, 1 script, 4 templates. v1.0.0 → v1.3.0 across 3 SkillOpt epochs. Signed-off-by: Magnus Hedemark <magnus919@pm.me>
4.7 KiB
4.7 KiB
Evaluation Workflow — ParamTuner and Metrics
This reference shows how to set up systematic evaluation for a LlamaIndex RAG pipeline, including parameter tuning, evaluator configuration, and batch scoring.
Setup
from llama_index.core.evaluation import (
FaithfulnessEvaluator,
RelevancyEvaluator,
SemanticSimilarityEvaluator,
BatchEvalRunner,
)
from llama_index.llms.openai import OpenAI
gpt4 = OpenAI(model="gpt-4o")
faith_evaluator = FaithfulnessEvaluator(llm=gpt4)
rel_evaluator = RelevancyEvaluator(llm=gpt4)
sim_evaluator = SemanticSimilarityEvaluator(llm=gpt4)
Single Query Evaluation
response = query_engine.query("What is the rate limit for the API?")
faith_result = faith_evaluator.evaluate_response(response=response)
print(f"Faithfulness: {faith_result.score} — {faith_result.feedback}")
# Example output: Faithfulness: 0.92 — The answer is grounded in the provided context
rel_result = rel_evaluator.evaluate_response(
response=response,
question="What is the rate limit for the API?"
)
print(f"Relevancy: {rel_result.score} — {rel_result.feedback}")
# Example output: Relevancy: 0.88 — The answer addresses the core question
Batch Evaluation
eval_questions = [
"How do I authenticate?",
"What are the rate limits?",
"How do I paginate results?",
"What error codes exist?",
"How do I handle webhooks?",
]
# Get responses
responses = [query_engine.query(q) for q in eval_questions]
# Batch evaluate
runner = BatchEvalRunner(
{
"faithfulness": FaithfulnessEvaluator(),
"relevancy": RelevancyEvaluator(),
},
workers=4,
)
results = runner.evaluate_responses(eval_questions, responses)
for metric_name, metric_results in results.items():
scores = [r.score for r in metric_results]
print(f"{metric_name}: mean={sum(scores)/len(scores):.2f}, "
f"min={min(scores):.2f}, max={max(scores):.2f}")
# Example output:
# faithfulness: mean=0.91, min=0.78, max=1.00
# relevancy: mean=0.85, min=0.72, max=0.94
ParamTuner — Systematic Optimization
from llama_index.core import VectorStoreIndex
from llama_index.core.param_tuner.base import ParamTuner
from llama_index.core.evaluation import SemanticSimilarityEvaluator
import numpy as np
def build_and_evaluate(params):
chunk_size = params["chunk_size"]
top_k = params["top_k"]
# Build index with these parameters
index = VectorStoreIndex.from_documents(
documents,
transformations=[SentenceSplitter(chunk_size=chunk_size)]
)
# Query
query_engine = index.as_query_engine(similarity_top_k=top_k)
responses = [query_engine.query(q) for q in eval_questions]
# Evaluate
evaluator = SemanticSimilarityEvaluator(llm=gpt4)
scores = []
for i, resp in enumerate(responses):
result = evaluator.evaluate_response(
response=resp,
reference=reference_answers[i]
)
scores.append(result.score)
return np.mean(scores)
param_tuner = ParamTuner(
param_fn=build_and_evaluate,
param_dict={
"chunk_size": [256, 512, 1024],
"top_k": [2, 5, 10],
},
fixed_param_dict={
"documents": documents,
"eval_questions": eval_questions[:3],
},
)
results = param_tuner.tune()
best = results.best_run_result
print(f"Best: chunk_size={best.params['chunk_size']}, "
f"top_k={best.params['top_k']}, score={best.score:.3f}")
# Example output: Best: chunk_size=512, top_k=5, score=0.894
Seven RAG Measurement Aspects
Based on (Gao, Yunfan et al., 2023), evaluate across these dimensions:
| Aspect | What it measures | Evaluator |
|---|---|---|
| Answer Relevance | Does the answer address the question? | RelevancyEvaluator |
| Context Relevance | Is the retrieved context on-topic? | RelevancyEvaluator (on context) |
| Faithfulness | Is the answer grounded in context? | FaithfulnessEvaluator |
| Context Recall | Are all needed chunks retrieved? | Custom (check coverage) |
| Context Precision | Are irrelevant chunks excluded? | Custom (check rank order) |
| Noise Sensitivity | Does noise degrade answers? | Compare with/without noise |
| Answer Correctness | Are facts correct? | SemanticSimilarityEvaluator |
Evaluation Best Practices
- Use held-out queries — never tune on the same questions you evaluate on
- Run evaluation in the same process — span-attached scoring preserves trace context
- Monitor in production — offline notebook scoring catches known issues; production monitoring catches novel ones
- Combine multiple metrics — faithfulness alone misses relevancy failures and vice versa
- Track over time — regressions are easier to catch when you have a baseline