Accident, defect and outage investigations end with a decision that ordinary question answering never faces: whether the evidence gathered so far is enough to close the case. We study this decision for LLM investigators, which request evidence from a case file, revise their hypotheses, and either close the case with a conclusion grounded in what they read or leave it open and name what is missing. This judgment does not come with capability: an untrained 9B model overstates its evidence in 97% of its answers, and a frontier model that identifies the right cause in 84% of cases still overstates in 91% and closes 17 of the 41 cases whose official finding is “cause undetermined”. Measuring it is also non-trivial: the source of a case largely predicts its label, and a rule that reads only the source reaches 83.0 balanced accuracy on our test cases. We therefore evaluate closure with three tests: closure accuracy, reported against this rule and within each source; evidence dependence, which removes the grounds of a conclusion and checks whether the model stops closing; and conclusion and gap quality, a judged checklist of what the model asserts and what it says is missing. We build Nautil, 731 audited cases from aviation, rail, maritime, chemical-safety and vehicle-defect reports and production server incidents, with teacher trajectories, an out-of-distribution test set and counterfactual evidence versions. Fine-tuning a 9B model on these trajectories makes its closures follow the evidence: removing the grounds lowers its closure rate by 26 points relative to a matched control (base model: 6; frontier model: 7), overstatement falls from 97% to 35%, and conclusions that are both correct and not overstated rise from 3% to 43% (frontier model: 9%). Reinforcement learning that rewards only the closure decision then raises balanced accuracy from 69.2 to 83.3, on par with the teacher, and within-source accuracy from 60.4 to 74.1, at some cost in evidence dependence. We release the models, the dataset and the evaluation suite.
@article{bi2026nautil,title={Not Until the Evidence Says So: Teaching LLM Investigators When to Close a Case},author={Bi, Tingzhu and Wang, Ping and Ma, Meng},journal={Preprint},year={2026},}
Preprint
JustDiag: A Diagnostic Justification Engine for Accountable Root Cause Analysis
Tingzhu Bi, Xinrui Jiang, Xun Zhang, and 5 more authors
Large language models can produce fluent root cause analyses, but fluent final answers alone are insufficient evidence for accountability in high-stakes operations. In real incident response, engineers need to know what evidence supported a diagnosis, which alternatives were considered, where contradictions remained, and whether the system resolved the case or preserved uncertainty. We address this gap with JustDiag, a diagnostic justification engine for RCA that maintains an explicit process state over evidence, findings, competing hypotheses, conflicts, and next checks. We evaluate the system on 66 real-world incidents using a two-layer protocol that separately scores final-answer quality and process quality. Relative to a matched control without diagnostic justification, JustDiag achieves stronger outcome and process scores, while accepting slightly lower terminal completion due to more calibrated non-closure. These results suggest that accountable RCA requires explicit diagnostic-justification artifacts and process-aware evaluation, not only fluent final answers.
@article{bi2026justdiag,title={JustDiag: A Diagnostic Justification Engine for Accountable Root Cause Analysis},author={Bi, Tingzhu and Jiang, Xinrui and Zhang, Xun and Su, Pengcheng and He, Congjie and Li, Jinglin and Wang, Ping and Ma, Meng},journal={arXiv preprint arXiv:2606.19407},year={2026},}
KDD
CAVIAR: Disentangling Root Causes with an ICA-based VAE for Large-Scale Microservice Systems
Xinrui Jiang, Tingzhu Bi, Meng Ma, and 1 more author
In Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD), 2026
CAVIAR disentangles the root causes of failures in large-scale microservice systems with an ICA-based variational autoencoder, separating independent causal factors from entangled performance signals to improve root-cause identification.
@inproceedings{jiang2026caviar,title={CAVIAR: Disentangling Root Causes with an ICA-based VAE for Large-Scale Microservice Systems},author={Jiang, Xinrui and Bi, Tingzhu and Ma, Meng and Wang, Ping},booktitle={Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD)},pages={520--531},year={2026},doi={10.1145/3770854.3780229}}
2025
NeurIPS
UnCLe: Towards Scalable Dynamic Causal Discovery in Non-linear Temporal Systems
Tingzhu Bi, Yicheng Pan, Xinrui Jiang, and 3 more authors
In Advances in Neural Information Processing Systems (NeurIPS), 2025
Uncovering cause-effect relationships from observational time series is fundamental to understanding complex systems. While many methods infer static causal graphs, real-world systems often exhibit dynamic causality where relationships evolve over time, requiring time-resolved causal graphs. We propose UnCLe, a deep learning method for scalable dynamic causal discovery. UnCLe employs a pair of Uncoupler and Recoupler networks to disentangle input time series into semantic representations and learns inter-variable dependencies via auto-regressive Dependency Matrices, estimating dynamic causal influences by analyzing datapoint-wise prediction errors induced by temporal perturbations. Extensive experiments show that UnCLe not only outperforms state-of-the-art baselines on static causal-discovery benchmarks but, more importantly, accurately captures evolving temporal causality in both synthetic and real-world dynamic systems (e.g., human motion).
@inproceedings{bi2025uncle,title={UnCLe: Towards Scalable Dynamic Causal Discovery in Non-linear Temporal Systems},author={Bi, Tingzhu and Pan, Yicheng and Jiang, Xinrui and Sun, Huize and Ma, Meng and Wang, Ping},booktitle={Advances in Neural Information Processing Systems (NeurIPS)},year={2025},}
2024
ICWS
G-Cause: Parameter-free Global Diagnosis for Hyperscale Web Service Infrastructures
Xinrui Jiang, Yang Zhang, Tingzhu Bi, and 8 more authors
In IEEE International Conference on Web Services (ICWS), 2024
G-Cause performs parameter-free global fault diagnosis across hyperscale web-service infrastructures, localizing root causes over large heterogeneous metric spaces without per-incident tuning.
@inproceedings{jiang2024gcause,title={G-Cause: Parameter-free Global Diagnosis for Hyperscale Web Service Infrastructures},author={Jiang, Xinrui and Zhang, Yang and Bi, Tingzhu and Shen, Xiangzhuang and Zhang, Yu and Pan, Yicheng and Ma, Meng and Han, Linlin and Wang, Feng and Liu, Xian and Wang, Ping},booktitle={IEEE International Conference on Web Services (ICWS)},pages={1003--1014},year={2024},doi={10.1109/ICWS62655.2024.00119}}
KDD
FaultInsight: Interpreting Hyperscale Data Center Host Faults
Tingzhu Bi, Yang Zhang, Yicheng Pan, and 7 more authors
In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD), 2024
Operating hyperscale data centers with millions of hosts makes fault diagnosis extremely intricate. Prior state-of-the-art methods use time-series causal discovery over homogeneous service-level metrics but fail on heterogeneous host-level metrics. We present FaultInsight, a highly interpretable deep causal host-fault diagnosis framework that offers diagnostic insights from multiple perspectives to reduce human troubleshooting effort. Evaluated on dozens of incidents from a production environment, FaultInsight delivers markedly better root-cause identification accuracy than SOTA baselines and shows strong deployability, helping engineers quickly understand the mechanisms behind faults.
@inproceedings{bi2024faultinsight,title={FaultInsight: Interpreting Hyperscale Data Center Host Faults},author={Bi, Tingzhu and Zhang, Yang and Pan, Yicheng and Zhang, Yu and Ma, Meng and Jiang, Xinrui and Han, Linlin and Wang, Feng and Liu, Xian and Wang, Ping},booktitle={Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD)},pages={141--152},year={2024},doi={10.1145/3637528.3672051},}
2023
SIGMETRICS
HEAL: Performance Troubleshooting Deep inside Data Center Hosts
Yicheng Pan, Yang Zhang, Tingzhu Bi, and 8 more authors
Proceedings of the ACM on Measurement and Analysis of Computing Systems (POMACS / ACM SIGMETRICS), 2023
HEAL troubleshoots performance problems deep inside data-center hosts by reasoning over heterogeneous host-level signals to localize the components responsible for degradation.
@article{pan2023heal,title={HEAL: Performance Troubleshooting Deep inside Data Center Hosts},author={Pan, Yicheng and Zhang, Yang and Bi, Tingzhu and Han, Linlin and Zhang, Yu and Ma, Meng and Shen, Xiangzhuang and Jiang, Xinrui and Wang, Feng and Liu, Xian and Wang, Ping},journal={Proceedings of the ACM on Measurement and Analysis of Computing Systems (POMACS / ACM SIGMETRICS)},volume={7},number={3},pages={54:1--54:24},year={2023},doi={10.1145/3626785}}
2022
ISSRE
VECROsim: A Versatile Metric-oriented Microservice Fault Simulation System
Tingzhu Bi, Yicheng Pan, Xinrui Jiang, and 2 more authors
In IEEE International Symposium on Software Reliability Engineering (ISSRE), Tools and Artifact Track, 2022
Most incidents in commercial cloud systems are not publicly available, forcing researchers to build ad-hoc experimental systems that cannot easily refactor functionality, scale architecture, or customize fault characteristics. We develop VECROsim, a versatile metric-oriented microservice fault simulation system, and release the VECROsim benchmark dataset. VECROsim is a highly customizable toolkit that automatically generates abnormal performance-metric datasets for microservice systems on demand, providing a standardized data foundation for fault-diagnosis and causal research.
@inproceedings{bi2022vecrosim,title={VECROsim: A Versatile Metric-oriented Microservice Fault Simulation System},author={Bi, Tingzhu and Pan, Yicheng and Jiang, Xinrui and Ma, Meng and Wang, Ping},booktitle={IEEE International Symposium on Software Reliability Engineering (ISSRE), Tools and Artifact Track},pages={297--308},year={2022},doi={10.1109/ISSRE55969.2022.00037},}