A Jev-driven pipeline selects investigation targets and organizes evidence to support diagnoses. In a recent study, researchers Yiming Su, Saad Mohammad Rafid Pial, Jackson Clark, and Tianyin Xu tested Jev as a decision aid for diagnosing and repairing incidents without an LLM agent. The pipeline collects and organizes cluster evidence, allowing Jev to identify likely root causes and assemble diagnosis reports.
In testing across 21 SREGym-Lite faults, the Jev-driven pipeline achieved a success rate of 76.2%, with a median diagnosis time of 14.6 seconds. Jev operates by choosing from a set of supplied options based on the evidence presented. The pipeline includes a collector that reads Kubernetes objects, events, recent pod logs, and resource usage, summarizing observations by component.
For example, in the case of the nginx-thrift fault, the collector identified a memory limit mismatch caused by a mutating admission webhook. Jev used this information to select the most relevant evidence and submit a diagnosis. The study found consistent results across multiple runs, with either all attempts passing or failing for each fault.
The researchers noted that the granularity of the cluster state presented to Jev is crucial; too coarse a summary may obscure important details, while excessive detail can overwhelm the model. Jev's performance was comparable to that of GPT-5.6 Sol, running significantly faster and at a lower cost.
However, the study identified two failure modes: Jev sometimes selected incorrect clues or lacked decisive evidence. Future work will focus on extending the pipeline to handle faults that span multiple services and involve evolving metrics, potentially integrating smaller models to enhance diagnosis capabilities. The findings suggest that incorporating Jev-driven diagnosis into SRE workflows could improve efficiency in identifying and resolving incidents.