You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
A curated list of research on LLM-agent traces: evidence tracing, execution provenance, failure attribution, observability, runtime safety, memory provenance, and learning from traces. Companion list of arXiv:2606.04990.
An evidence-annotated benchmark for RAG over university course syllabi — 80 questions with document, page and supporting-span labels, including unanswerable items, evaluated across closed-book, full-context, basic RAG and citation-constrained RAG.