JCTAM OPEN ACCESS

Journal of Computer Technology and Applied Mathematics

ISSN:3007-4126 (print) | ISSN:3007-4134 (online) | Publication Frequency: Bimonthly

OPEN ACCESS|Research Article||31 August 2026

Tabular In-Context Learning for Low-Resource Rare Industrial Fault Detection: a Cost-Sensitive Comparison on Scania APS and SECOM

* Corresponding Author1: Guoqing Song, E-Mail: m19896144341@163.com

Publication

Accepted 2026 August 7 ; Published 2026 August 31

Journal of Computer Technology and Applied Mathematics, 2026, 3(4), 3007-4126.

Abstract

Rare industrial fault detection is complicated by class imbalance, missing sensor values, asymmetric errors, and limited training resources. We compared TabICLv2 with TabM, XGBoost, and logistic regression on Scania APS Failure and SECOM. For Scania, nested stratified samples of 5,000 and 10,000 training records were drawn under five prespecified seeds. Shared three-fold splits supported out-of-fold (OOF) evaluation; tunable models were selected by OOF AUPRC, and decision thresholds were locked by minimizing OOF cost, defined as C = 10 × FP + 500 × FN. The official test set remained locked until all selections were complete. SECOM used three repeats of five-fold outer cross-validation with three-fold inner selection. Paired class-stratified bootstrap analyses used 2,000 replicates. At the prespecified 10,000-record Scania budget, TabICLv2 achieved an AUPRC of 0.8879, an official cost of 13,002, an MCC of 0.6514, and a Brier score of 0.00685; XGBoost achieved 0.8549, 16,456, 0.6125, and 0.00840, respectively. The TabICLv2-minus-XGBoost difference was 0.03294 for AUPRC (95% CI 0.02089–0.04570) and −3,454 for cost (95% CI −5,370.60 to −1,711.95). Across 15 SECOM outer folds, mean AUPRC was 0.2083 for TabICLv2 and 0.1705 for XGBoost. After averaging repeated OOF probabilities per record, the paired AUPRC difference was 0.05110 (95% CI −0.00693 to 0.11358). TabICLv2 required 992.2 s, approximately 413 times the XGBoost time. TabICLv2 improved the prespecified Scania outcomes but incurred substantial task-side computational cost; SECOM provided suggestive rather than definitive support.

Keywords

Tabular Foundation Model , In-context Learning , Industrial Fault Detection , Class Imbalance , Cost-sensitive Learning .

Metadata

1-9

18

Software Systems

Software Engineering

Cite This Article

APA Style

Song, G. (2026). Tabular in-context learning for low-resource rare industrial fault detection: a cost-sensitive comparison on scania aps and secom. Journal of Computer Technology and Applied Mathematics, 3(4), 1-9. https://doi.org/10.70393/6a6374616d.343334

Acknowledgments

Not Applicable.

FUNDING

Not Applicable.

INSTITUTIONAL REVIEW BOARD STATEMENT

Not Applicable.

DATA AVAILABILITY STATEMENT

Not Applicable.

INFORMED CONSENT STATEMENT

Not Applicable.

CONFLICT OF INTEREST

Not Applicable.

AUTHOR CONTRIBUTIONS

Not application.

References

1.
He, H., & Garcia, E. A. (2009). Learning from imbalanced data. IEEE Transactions on knowledge and data engineering, 21(9), 1263-1284.

2.
Saito, T., & Rehmsmeier, M. (2015). The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets. PloS one, 10(3), e0118432.

3.
Chicco, D., & Jurman, G. (2020). The advantages of the Matthews correlation coefficient (MCC) over F1 score and accuracy in binary classification evaluation. BMC genomics, 21(1), 6.

4.
Glenn, W. B. (1950). Verification of forecasts expressed in terms of probability. Monthly weather review, 78(1), 1-3.

5.
Fouladvand, S., Noshad, M., Goldstein, M. K., Periyakoil, V. J., & Chen, J. H. (2023). Mild cognitive impairment: data-driven prediction, risk factors, and workup. AMIA Summits on Translational Science Proceedings, 2023, 167.

6.
Grinsztajn, L., Oyallon, E., & Varoquaux, G. (2022). Why do tree-based models still outperform deep learning on typical tabular data?. Advances in neural information processing systems, 35, 507-520.

7.
Ye, H. J., Liu, S. Y., Cai, H. R., Zhou, Q. L., & Zhan, D. C. (2024). A closer look at deep learning methods on tabular datasets. arXiv preprint arXiv:2407.00956.

8.
Hollmann, N., Müller, S., Purucker, L., Krishnakumar, A., Körfer, M., Hoo, S. B., ... & Hutter, F. (2025). Accurate predictions on small data with a tabular foundation model. Nature, 637(8045), 319-326.

9.
Qu, J., HolzmÞller, D., Varoquaux, G., & Morvan, M. L. (2025). Tabicl: A tabular foundation model for in-context learning on large data. arXiv preprint arXiv:2502.05564.

10.
Qu, J., HolzmÞller, D., Varoquaux, G., & Morvan, M. L. (2026). TabICLv2: A better, faster, scalable, and open tabular foundation model. arXiv preprint arXiv:2602.11139.

11.
Gorishniy, Y., Kotelnikov, A., & Babenko, A. (2025, May). Tabm: Advancing tabular deep learning with parameter-efficient ensembling. In International Conference on Learning Representations (Vol. 2025, pp. 77899-77935).

12.
Niculescu-Mizil, A., & Caruana, R. (2005, August). Predicting good probabilities with supervised learning. In Proceedings of the 22nd international conference on Machine learning (pp. 625-632).

13.
Varma, S., & Simon, R. (2006). Bias in error estimation when using cross-validation for model selection. BMC bioinformatics, 7(1), 91.

14.
Cawley, G. C., & Talbot, N. L. (2010). On over-fitting in model selection and subsequent selection bias in performance evaluation. The Journal of Machine Learning Research, 11, 2079-2107.

15.
Scania CV AB. (2017). APS failure at Scania Trucks [Data set]. UCI Machine Learning Repository. https://doi.org/10.24432/C51S51

16.
McCann, M., & Johnston, A. (2008). SECOM [Data set]. UCI Machine Learning Repository. https://doi.org/10.24432/C54305

17.
Effron, B., & Tibshirani, R. J. (1993). An introduction to the bootstrap. Monographs on statistics and applied probability, 57, 436.

18.
Loza, M., Chushig-Muzo, D., Milara, E., Bote-Curiel, L., Estrada-Petrocelli, L., & Grijalva, F. (2026). Empirical Evaluation of Out-Of-Distribution Performance of Tabular Foundation Models. arXiv preprint arXiv:2607.26000.

PUBLISHER'S NOTE

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.

cc Copyright © 2025 The Author(s). Published by Southern United Academy of Sciences.
This work is licensed under a Creative Commons Attribution 4.0 International License , which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.