A question about software operations

Can a language model connect the symptoms of a software failure to its underlying cause? OpenRCA places this question in a software operations setting, where a natural-language request must be answered using telemetry rather than the prompt alone.

Reasoning across evidence

The benchmark brings together KPI time series, dependency traces, and logs. Its RCA-agent baseline uses Python for data retrieval and analysis, allowing the model to work with substantial telemetry without putting every record into its context.

Explore the work

The official repository provides the benchmark, baseline, and reproduction instructions. The associated paper appeared at ICLR 2025.