0% found this document useful (0 votes)
10 views15 pages

Financial Risk Prediction for SMEs in Vietnam

This research utilizes the TabNet deep learning model to predict financial risk for Vietnamese listed SMEs, analyzing 1,181 quarterly observations from 34 firms between 2016 and 2024. The findings indicate that leverage and debt structure are critical indicators of financial distress, with TabNet achieving a balanced accuracy of 78.2%, surpassing traditional methods like logistic regression. The study emphasizes the importance of optimizing capital structures for SMEs in Vietnam's bank-dominated financial landscape, providing insights for stakeholders and enhancing predictive models for financial distress in emerging markets.

Uploaded by

quangnd
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOC, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
10 views15 pages

Financial Risk Prediction for SMEs in Vietnam

This research utilizes the TabNet deep learning model to predict financial risk for Vietnamese listed SMEs, analyzing 1,181 quarterly observations from 34 firms between 2016 and 2024. The findings indicate that leverage and debt structure are critical indicators of financial distress, with TabNet achieving a balanced accuracy of 78.2%, surpassing traditional methods like logistic regression. The study emphasizes the importance of optimizing capital structures for SMEs in Vietnam's bank-dominated financial landscape, providing insights for stakeholders and enhancing predictive models for financial distress in emerging markets.

Uploaded by

quangnd
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOC, PDF, TXT or read online on Scribd

FINANCIAL RISK PREDICTION WITH TABNET: THE

PREDOMINANT ROLE OF DEBT STRUCTURE IN LISTED SMALL


AND MEDIUM ENTERPRISES IN VIETNAM
Duy Quang, Nguyen1

Abstract:
The research uses TabNet, a deep learning model based on attention mechanisms,
and predicts financial risk for Vietnamese listed SMEs by conducting 1,181
quarterly observations on 34 firms at the Hanoi Stock Exchange (2016–2024)
period. The results show that leverage and debt structure metrics, particularly the
nonlinear threshold effect of the debt-to-equity ratio, are the most powerful
indicators of financial distress. TabNet is found to outperform the traditional
approaches significantly with a 78.2% balanced accuracy rate as opposed to
logistic regression's 69.4%, and it also provides better interpretability through its
sequential attention mechanism. The model exposes a decision-making hierarchy
which gives priority to capital structure first, then profitability, and liquidity and
cash flow last. This study not only contributes to the development of predictive
models for financial distress in emerging markets, but also provides practical
recommendations for optimizing SME capital structures in financially dominated
by banks systems.

Keywords: Financial Forecast; Financial Risk; Machine learning interpretability;


SMEs; Vietnam
JEL: G170; G320

1. Introduction

A forecast of financial risk is vital for stability and growth in the economy, especially for
emerging markets in which small and medium-sized enterprises (SMEs) constitute a
significant landscape. While mathematical models are considered strong tools to assess
financial vulnerability, the prediction of distress still remains complex in the case of SMEs,
specifically under dynamic settings with sudden economic changes (Krüger & Meyer, 2021).
In Vietnam, SMEs listed on Hanoi Stock Exchange (HNX) represent a sizable workforce
segment in highly valuable sectors and also pay large chunks of tax revenue to the nation
(Diep, 2024), thus highlighting the need for such entities and their financial health
assessment. Their stability influences both their shareholders and the wider realization of
employment, supply chains, and regional economic development (Thu & Xuan, 2023).
Contemporary approaches to financial distress prediction for SMEs in emerging markets have
many shortcomings. The first limitation relates to a persistent trade-off between model
complexity and interpretability; highly accurate models have always lacked the transparency
needed for intervention weighing; being able to recognize the causal mechanisms is as

1
Duy Quang Nguyen, Head of Digital Economics Major, Faculty of Economics, Ho Chi Minh City University of Economics
and Finance, Tel : +84 905 083 366; Email: quangnd@[Link]

1
important as the actual prediction of distress (Gao et al., 2025). Another trap in the race:
financial distress models are prepared in the setting of large entities in the developed
economies and fail to capture peculiarities of SMEs in emerging markets vis-a-vis differences
in capital structures, information environments, and patterns of distress (Leipziger et al.,
2024). Third, the static approach is another limitation of current frameworks, since it looks at
a sequential development of risk, intoxication from distress signals on financial statements
(Yu et al., 2025). Lastly, there still is a considerable implementation gap between theoretically
advanced models and practical implementations in resource-restrained environments (Dodd et
al., 2020).
The study fills those gaps using TabNet, an attention-based neural network architecture, for
financial-risk prediction for listed SMEs in Vietnam. TabNet offers high predictive accuracy
with its attention mechanism and enhanced interpretability-limiting similarity to financial
analysts focusing on the relevant variables for risk assessment. This methodology thus
facilitates the correct prediction of financial distress and the identification of primary risk
drivers, enriching both theory and practice of financial-risk management.
Specifically, this study looks into: (1) how well TabNet can classify financial risk with
interpretability and whether the attention mechanism could help us grasp nonlinear
relationships and ascertain some insights into feature importance; (2) the importance levels of
indicators, that is, financial ratios in risk determination and how the ranking compares with
current theoretical and empirical literature; (3) if temporal patterns exhibited by financial
indicators provide an early warning signal for distress and which among them show the
strongest lead times; (4) how debt structure and cash flow indicators affect risk classification
mostly when considered alongside traditional financial ratios; and (5) how the results of risk
classification are affected by nonlinear relationships and interaction effects between financial
ratios.
Such findings expand the very core of financial risk modelling in emerging markets through
methodological innovations and empirical insight into the determinants of financial health in
Vietnamese SMEs. The results also carry practical implications for financial institutions,
regulators, investors, and SMEs, thus enabling more informed decision making and leading to
the stability of the economic ecosystem in Vietnam's burgeoning capital markets

2. Theoretical Background

2.1. Economic Theories of Financial Risk and Financial Distress


Different economic theories underlie interpretations of financial risk and provide guidance
with respect to selecting variables for the forecasting model. Capital Structure Theory (Kraus
& Litzenberger, 1973) postulates that firms adjust their capital structure to balance the tax
shield with the costs of financial distress. A higher degree of leverage presents more risk due
to an inability of cash flow to provide for debt servicing. This theory directly motivates our
viewing leverage ratios and debt structure indicators as key predictive variables.
The Pecking Order Theory (Myers & Majluf, 1984) shows the information asymmetry that
lies between firms and investors, such that firms prefer to use their internal funds before
resorting to debt and then equity. Given that SMEs arguably suffer the most from illiquid
access to external capital, this potentially raises their vulnerability since they have to rely on
depleted sources for funds. Agency Theory (Jensen & Meckling, 1976) focuses on conflicts
between shareholders and creditors in situations of financial distress. Shareholders may then
2
pursue high-risk courses of action that transfer wealth away from the creditors. This dynamic
is even more important for SMEs with some kind of leverage. Recent extensions of the
framework have examined governance mechanisms and ownership structure in emerging
markets (Kim, 2022), a setting that is directly related to Vietnam.
Per the proclamation of Fama (1970), an Efficient Market is semi-strong and implies that by
virtue of public availability, information would have caused price changes in securities. As a
matter of fact, most buyers and sellers would accept it to be true in developed markets, while
in an under-developed market such as Vietnam, it goes completely undebated. In these
instances, accounting ratios would be stronger forecasters than market-based indicators,
thereby hinting at the approach of variable selection (Cakici et al., 2021). Complexity Theory
(Arthur, 2021) brings to light that nonlinear interactions and feedback loops can occur within
firms arising out of so-called complex adaptive systems; for this reason, one can argue the
relevance of machine learning approaches, like the one employed by TabNet, which are
capable of unearthing relationships that escape identification through traditional linear
methods. While Institutional Theory (Kakooza et al., 2024) highlights that an appropriate
consideration of the regulatory environment remains of utmost importance in identifying the
financial risk relationships, it should be acknowledged that as a transitional economy,
Vietnam would carry peculiar distress dynamics that would necessitate a relevant
consideration of the institutional framework (Cakici et al., 2021).
2.2. Mathematical Models for Financial Risk Forecasting
TabNet is the most recent development in the long and interesting line of financial risk
modeling. Initial methods like Altman’s (1968) multivariate discriminant analysis (MDA)
gave the basis for distress prediction. But MDA makes certain assumptions, such as
multivariate normality and equal variance for all groups, which are often not the case. On the
other hand, logistic regression interprets financial risk in probabilistic terms, it loosens up
some MDA constraints, and takes the linearity in log odds still as an assumption. Survival
Analysis is the one that makes the time to financial distress a variable and thus captures the
gradual change of risk (Ridzwan et al., 2024).
The traditional statistical models have many drawbacks, and the main drawback is their
inability to model complex nonlinear relationships between factors and variable interactions.
Therefore, the need for machine learning has risen. On the other hand, tree-based models are
much more flexible but do lack in transparency in most cases. In contrast to general of neural
networks, which provide powerful alternatives for the upcoming closer modeling with the
ability to approximate highly complex nonlinear functions. The choice of TabNet is justified
by its interpretable nature through the sparsemax attention mechanism. This is absolutely
opposite to the hidden nature of the majority of the neural network architectures.
Interpretability provided is vital for characterization of the distress among Vietnamese SMEs.
2.3. TabNet and its Sparsemax Attention Mechanism
The architecture of TabNet is known for its sequential attention mechanism, which, in a way,
imitates the decision-making processes of financial analysts (Table 1). In contrast to
conventional neural networks, which deal with all of the input features at once, TabNet works
with them one at a time through the use of Sparsemax that generates at each decision step
sparse feature selection matrices. This feature of sparsity not only contributes to the
interpretability of the model but also makes it possible to see which features are the most
influential at each step.
3
Table 1. Comparison of Softmax vs. Sparsemax Attention Mechanisms
Feature Softmax Sparsemax
Formula
Sparsity
Gradient
Computationa
l Complexity
Intuitive Probability distribution based on exponential Euclidean projection of z onto the probability
Explanation normalization simplex
Source: Author’s Summary

The sparsemax function maps attention scores to the probability simplex, thus producing
sparse probability vectors that comprise a lot of zeros, whereas softmax usually gives nonzero
probabilities to all features. The sparse feature selection matrix created in this way is suitable
for processing in a more efficient way and model transparency is improved, hence it is
specially good for the financial indicators of SMEs to be analyzed in terms of their complex
interplay. Processing in sequence not only improves the interpretability of the model but also
shows the order of the features' contributions to the risk prediction.
2.4. Economic Interpretation of the Model
The financial distress probability indicator. Economic theories can be used to interpret these
coefficients. The Capital Structure Theory predictions get support from a positive coefficient
for leverage ratios, meaning financial distress probability goes up with more debt. Sequential
feature processing makes it possible to study the chronological dynamics of financial distress,
which is consistent with the fact that financial weaknesses are gradually developing.
Nonlinear relationships capturing ability of the model goes beyond traditional models
limitations and gives a detailed understanding of the economic mechanisms that cause SMEs
financial risk.
2.5. Hypotheses
The unified theoretical framework leads the author to propose the following hypotheses:
H1: The Influence of Leverage and Debt Structure: In line with Capital Structure Theory
(Kraus & Litzenberger, 1973), accounting for leverage and debt structure is the main predictor
of financial distress for listed Vietnamese SMEs. This assumption is also indicative of limited
capital accessibility for SMEs and hence, the bank-centric financing in Vietnam
(Boyarchenko & Elias, 2024; Boston, 2020).
H2: Temporal sequence of financial distress indicators: The predictive significance of
financial indicators has a temporal sequence, wherein profitability and operational efficiency
indicators are the first to raise alarms and liquidity indicators the last. This builds upon the
conclusions of Freiesleben et al. (2024) and utilizes the attention mechanism of TabNet to
discern the order in which the indicators support the prediction of distress.
H3: Superiority of TabNet in Capturing Nonlinear Financial Dependencies: TabNet is on
top of traditional statistical methods when it comes to predicting the financial distress as it is
capable of unveiling complex non-linear relationships among financial variables. To put it
differently, it is in line with Complexity Theory (Arthur, 2021) and also has the backing of
earlier research that has proved TabNet’s outstanding performance in like financial prediction
contexts (Mai et al., 2019).
4
H4: The Predominant Role of Short-term Debt: The indicators of short-term debt will be
found to be more critical than the long-term debt indicators in predicting financial distress.
This is based on the poor access of SMEs to long-term capital; hence, they are most likely to
suffer from the risks of rollover and interest rate volatility (Gupta et al., 2018).
H5: Added Value of Cash Flow Indicators: Cash flow indicators are very useful as they
point out the extra predictive information, especially in the case of identifying risks among
financially stable firms that rely solely on accrual accounting indicators for assessment. This
signals the importance of cash flow in situations where the transparency of accounting may
not be very good (Seretidou et al., 2025).
The presented hypotheses form a broad empirical analysis framework by amalgamating
economic theories with TabNet's skills to fill the identified research gaps.

3. Research Methodology

The research predicts the financial risk of small and medium-scale listed companies present
on the Hanoi Stock Exchange (HNX) using a numerical method and financial ratio
examination together with the TabNet machine learning model. The performance of the
TabNet model has been the major reason behind its selection, which can learn to describe
nonlinear relationships in financial data, while still being more interpretable than the
traditional "black box models" (Arik & Pfister, 2021). This interpretability provides
actionable insights for the stakeholders and at the same time confirms the existing economic
theories regarding the firm's distress.
Our research work has used a cross-sectional dataset along with lagged variables in order to
reflect the very nature of the financial risk capturing its dynamic aspect. The data for our
study consisted of the 34 small and medium-sized listed companies (having capital less than
VND 100 billion) on HNX, and was obtained through quarterly observations from 2016 to
2024, leading to 1181 observations. This time period not only covers the pre- and post-
COVID-19 periods but also makes it possible to perform a robust model evaluation through
different economic conditions. The dependent variable is made up of a three-point ordinal
financial risk classification (low, medium, and high) based on a modified Altman Z-score
Altman et al. (2020) adapted for emerging market SMEs, which is more sophisticated than
two-classification and thus provides valuable insights into the financial distress "gray zone".
The independent variables used in the analysis are financial ratios that have been categorized
according to corporate finance theories: leverage and capital structure (Debt/Equity (DE),
Short-Term Debt to Total Debt (STDTD)); liquidity (Current Ratio (CR), Working Capital to
Total Assets (WCTA)); profitability (Return on Assets (ROA), Gross Profit Margin (GPM),
Return on Equity (ROE)); asset composition (Fixed Assets to Total Assets (FATA)); debt
coverage (Interest Coverage Ratio (ICR)); revenue growth (REVG); and cash flow (Operating
Cash Flow to Total Assets (OCFTA), Free Cash Flow to Total Assets (FCFTA)). To reflect
the temporal dynamics within the study, lagged variables for one year (L1) and two years (L2)
along with year-over-year change variables (Δ) for the key ratios were included. Power
analysis indicated that our sample was sufficiently large for detection of medium-to-large
effects (f² ≥ 0.15) at standard significance levels (α = 0.05, 1-β = 0.8).
The venues for data were Cafef and HNX. The preprocessing steps included defining &
harmonizing accounting terms, treating outliers by winsorizing the top (99th) and bottom (1st)
percentile, and missing-value handling via trend interpolation and industry average
5
substitution. All financial variables were standardized. To deal with the class imbalance (34%
low risk, 49% medium risk, and 17% high risk), the Synthetic Minority Oversampling
Technique (SMOTE) was used in the training set until all risk classes were equally
represented, without any bias being introduced. The dataset was prepared through stratified
sampling with an 80% training and a 20% testing set to maintain the same risk class
representation.
3.1. TabNet Model Specification
The architecture of TabNet included four decision steps (N=4), 64 as the feature
width, 0.2 as the dropout rate, and 1e-4 as the L2 regularization coefficient. These
hyperparameters were found by performing a grid search with 5-fold stratified cross-
validation, where the weighted average F1-score across the risk classes was maximized (refer
to Table 2). The attention component of the model allows for the selection of features in a
sequential manner and improves the understanding of the process by showing what features
were important at each stage of the decision. The Adam optimizer was employed for training,
starting with a learning rate of 0.01, which was reduced by a factor of 0.5 according to a
learning rate scheduler. The target function was a multiclass cross-entropy loss function with
class weights to deal with class imbalance.
Table 2. Hyperparameter Optimization Results
d_model \ N 1 2 3 4 5

8 0.683 0.705 0.712 0.718 0.719


16 0.692 0.721 0.735 0.742 0.743
32 0.701 0.735 0.749 0.758 0.76
64 0.713 0.741 0.757 0.764 0.766
128 0.708 0.736 0.753 0.761 0.759

Source: Author’s Analysis results

3.2. Feature Engineering and Selection


Our feature engineering framework is grounded in corporate finance theory, which
categorises variables into conceptually meaningful groups. Multicollinearity was addressed
through correlation matrix analysis, Variance Inflation Factor (VIF) calculations, and
Principal Component Analysis (PCA) where necessary. Feature selection was guided by both
theoretical relevance and empirical performance, ensuring that only the most informative
variables were included in the final model.
3.3. Comparison Models and Model Evaluation
In order to assess the performance of TabNet, we performed a comparison with widely
accepted machine learning models: Multinomial Logistic Regression, Random Forest,
XGBoost, and an ensemble strategy that included predictions from TabNet, Random Forest,
and XGBoost. Fair comparison was guaranteed as all models were trained and evaluated on
the same datasets using the same preprocessing techniques.
The performance of the model was strictly evaluated through metrics that are suitable for
imbalanced classification: balanced accuracy; weighted F1-score; Cohen's Kappa; confusion
matrices with precision, recall, and specificity calculations for each risk class; AUC-ROC;
ranked class error analysis; and feature importance analysis. So these metrics will give us a
very strong measure of how well the prediction performed, taking into account both the
6
correctness and the different types of misclassification costs (for instance, if a high-risk firm
is misclassified as low risk, the consequences will be greater than the reverse case).
3.4. Methodological Challenges and Mitigation Strategies
The methodological challenges which hindered the result's robustness were solved with
various techniques. The main method employed for dealing with the small sample size was
the application of k-fold stratified cross-validation (k=5), L2 regularisation, dropout, careful
feature selection, and rigorous hyperparameter tuning. The imbalance of the classes was dealt
with by SMOTE (Synthetic Minority Over-sampling Technique), the use of class weights in
the loss function, and the selection of the right evaluation metrics. Multicollinearity was
controlled by using correlation analysis, calculating variance inflation factor (VIF), applying
PCA, and using the feature selection capabilities inherent to TabNet. Temporal dependencies
were accounted for by employing lagged and change variables. The missing data were treated
by applying context-specific imputation and trend interpolation. Outliers were dealt with
firstly by windsorisation and secondly through data standardisation. The model
interpretability was enhanced through visualisation, feature importance analysis and partial
dependence plots. This comprehensive methodological approach not only provides strong,
trustworthy and interpretable results but also solves the specific issues of predicting financial
risk in the Vietnamese SME context.

4. Research Results and Discussion

4.1. Descriptive Statistics and Sample Characteristics


Our dataset consists of 1181 records coming from 34 small and medium-sized enterprises
(SMEs) that were traded on the Hanoi Stock Exchange during the period of 2016 to 2024.
Descriptive statistics for the significant financial variables that have been divided into
categories of leverage, liquidity, profitability, and cash flow metrics are communicated in
Table 3. The sample shows considerable diversity in terms of all financial aspects. The
average debt-to-equity ratio (DE) of 1.74 signals high leverage, and some companies have
gone so far as to exceed 7, which indicates a really risky financial situation. The short-term
debt (STDTD average: 0.83) relationship suggests that the company is mostly using short-
term debt which might lead to liquidity shock vulnerability.
The indicators of profitability reflect a small rise together with big oscillations, as ROA and
ROE are at an average of 2.8% and 6.8% respectively. The liquidity ratios (CR and WCTA)
suggest that, on the whole, the firms have enough liquidity but the lowest figures indicate
extreme difficulties for some of them. The interest coverage ratio (ICR) at 3.27 on average
indicates that the majority of firms have earnings enough to cover interest; however, the
lowest values (-2.83) indicate that in some cases there is very severe financial distress. The
cash flow metrics (OCFTA and FCFTA) show positive averages but have very high volatility
with a considerable number of negative values.
The Z-score which is 5.01 on average suggests that there is a moderate risk in total, however,
the mask hid that there were very different risk profiles in the sample. The distribution of risk
classifications reveals that 34% of companies fall under the safe category
(RISK_CATEGORY=0), 49% under the gray category (RISK_CATEGORY=1), and 17%
under the high-risk category (RISK_CATEGORY=2). The relatively small share of high-risk
companies may be an indication of the effect of HNX listing requirements, which do not
allow the most financially unstable companies to operate.
7
Table 3. Descriptive Statistics of Key Financial Variables
Variable Mean Median Standard Deviation Min Max
DE (Debt/Equity) 1.74 1.38 1.45 0.14 7.32
STDTD (Short-term Debt/Total Debt) 0.83 0.89 0.19 0.28 1
CR (Current Ratio) 1.53 1.22 1.01 0.42 5.18
WCTA (Working Capital/Total Assets) 0.12 0.09 0.23 -0.41 0.62
FATA (Fixed Assets/Total Assets) 0.34 0.31 0.22 0.02 0.82
ROA (Profit/Total Assets) 0.028 0.031 0.058 -0.132 0.176
GPM (Gross Profit Margin) 0.14 0.12 0.11 -0.06 0.49
ROE (Profit/Equity) 0.068 0.076 0.142 -0.384 0.401
ICR (Interest Coverage Ratio) 3.27 2.21 3.94 -2.83 18.65
REVG (Revenue Growth) 0.082 0.065 0.321 -0.582 1.204
OCFTA (CFO/Total Assets) 0.059 0.046 0.142 -0.315 0.412
FCFTA (FCF/Total Assets) 0.024 0.017 0.149 -0.384 0.376
Z-Score 5.01 4.87 1.83 1.24 9.38

Source: Author’s Analysis results

4.2. Statistical Distribution and Relationships Between Variables


The advanced statistical properties such as skewness, kurtosis, and results of the Jarque-Bera
and Augmented Dickey-Fuller (ADF) tests are shown in Table 4. All main variables have
their normality strongly rejected by the Jarque-Bera test which means that the distributions
are skewed and with heavy tails. Non-normality supports the use of nonlinear models like
TabNet. The ADF test reveals that the majority of the variables are stationary, except for ICR
which shows slight nonstationarity. Granger causality tests support the financial theory and
show that DE has a significant impact on both ROA and CR, while ROA affects OCFTA and
ICR.
The correlation analysis disclosed a strong negative correlation of -0.684 between DE and Z-
score, thus confirming the connection of high-leverage with increased financial risk. A
significant multicollinearity is observed among the pairs of the variables such as ROA and
ROE (0.912), OCFTA and FCFTA (0.912), and WCTA and CR (0.786). This potential
multicollinearity was solved by employing the feature-selection capabilities inherent in
TabNet. ROA shows stronger correlation to Z-score (0.478) than ROE does (0.324),
indicating that ROA is a stronger financial risk predictor and it is less influenced by leverage.
The cash flow indicators have moderate correlations with the Z-score which is less than that
of ROA, implying that the accrual-based accounting measures could be the better ones in
capturing the SMEs’ financial risk.
The scatter plot and cluster analyses also pointed to non-linear relationships among leverage,
profitability, liquidity, and risk, thus confirming the appropriateness of TabNet in modeling
these intricate interplays. There are three distinct clusters: Stable, Balanced, and Stressed,
with each showing different financial characteristics and risk profiles. The time-series
analysis emphasized the financial distress signs that had been warned of earlier such as a
falling ROA, a lowering CR and WCTA, a rise in DE, a worsening ICR, and a negative
OCFTA.
Table 4. Advanced Statistical Characteristics
Variable Skewness Kurtosis Jarque-Bera (p-value) ADF test (p-value) Granger causality

8
DE 1.73 4.82 <0.001 0.042 → ROA (0.003), → CR (0.015)
ROA -0.56 5.21 <0.001 0.008 → OCFTA (0.001), → ICR (0.012)
CR 1.84 7.26 <0.001 0.032 → WCTA (0.001)
WCTA 0.15 3.48 0.038 0.027 → ROA (0.089)
OCFTA 0.32 4.95 <0.001 0.004 → FCFTA (0.001), → ICR (0.045)
ICR 2.63 12.84 <0.001 0.073 There is no significant causal relationship

Source: Author’s Analysis results

4.3. TabNet Model Performance and Feature Importance


TabNet proved to be a very strong classifier in all the different risk categories. The metrics of
balanced accuracy (0.782) and Cohen's kappa (0.693) showed a significant advantage over
random classification (see Table 5). The model reached high F1 scores for all risk groups,
where slightly lower performance for the high-risk category was noticed, likely due to class
imbalance even though SMOTE was applied. An average AUC-ROC of 0.881 proved to be
excellent in terms of the ability to discriminate among different groups, especially the high-
risk group (0.907).
Table 5. TabNet Performance Metrics
Metric Value Standard Deviation (5-fold CV)
Balanced Accuracy 0.782 0.031
Cohen's Kappa 0.693 0.042
Weighted Average F1-Score 0.764 0.038
Average AUC-ROC 0.881 0.027
F1-Score (Low Risk) 0.812 0.043
F1-Score (Medium Risk) 0.776 0.035
F1-Score (High Risk) 0.702 0.057
Precision (Low Risk) 0.784 0.046
Precision (Medium Risk) 0.729 0.039
Precision (High Risk) 0.75 0.068
Recall (Low Risk) 0.846 0.039
Recall (Medium Risk) 0.827 0.034
Recall (High Risk) 0.654 0.072

Source: Author’s Analysis results

The application of bootstrap and Monte Carlo dropout methods for the analysis of model
uncertainty (Table 6) resulted in the identification of a high-risk group with a greater
uncertainty range, predominantly contributed by aleatoric uncertainty. The calibration curves
(Figure 1) reported a good calibration whilst the medium-risk class was an exception to this
claim.

Table 6. Model Uncertainty Analysis


Risk Average Predicted 95% Confidence CI Epistemic Aleatoric
Category Probability Interval Width Uncertainty Uncertainty
Low Risk 0.846 [0.782, 0.901] 0.119 0.032 0.087
Medium 0.827 [0.768, 0.886] 0.118 0.035 0.083
Risk
9
High Risk 0.654 [0.572, 0.736] 0.164 0.048 0.116

Source: Author’s Analysis results

Figure 1. Calibration Curves for Risk Classes

Source: Author’s Analysis results

Feature importance analysis (Table 7, Figure 2, Figure 3) provided evidence that DE was the
foremost predictor (17.2%) followed by ROA (11.8%) and ICR (9.7%). The attention
mechanism of TabNet revealed a series of decisions made one after the other, which initially
looked at leverage, profitability, and working capital and then shifted to debt repayment
capacity and liquidity. Eventually, it merged the cash flow indicators. The permutation feature
importance (Figure 10) supports the superiority of DE whereas the SHAP value analysis
(Figure 11) shows that capital structure, asset structure, and liquidity are the key aspects in
predicting risks.

Table 7. Feature Importance Analysis


Feature Step 1 Step 2 Step 3 Step 4
Mean Entropy 1.23 1.48 1.62 1.73
Sparsity (% features with weight > 0.01) 24.20% 31.50% 36.80% 42.10%
Number of features accounting for 80% of total weight 3 5 7 8
Correlation with previous step - 0.42 0.38 0.31
Most important feature DE ICR OCFTA FATA
Average proportion of the most important feature 36.40% 22.80% 25.30% 19.20%

Source: Author’s Analysis results

10
Figure 2. Importance of permutation features

Source: Author’s Analysis results

Figure 3. SHAP Value Analysis

Source: Author’s Analysis results

11
4.4. Comparative Analysis and Robustness Testing
TabNet is a remarkable model that has outperformed logistic regression in every evaluation
metric and this has been evidenced by the demonstration of the benefits of nonlinear modeling
inL the capture of complex relationships among the financial indicators and risk (see Table 9).
Besides, TabNet was able to outperform Random Forest and XGBoost by a small amount but
that difference was supported by the statistical evidence, which confirms his effectiveness in
the specific case mentioned. Moreover, combining all three methods in an ensemble model
resulted in the highest performance, which indicates that every single model contributed to the
data from its own and unique point of view. Besides, stability and generalisation of the
TabNet model across the different data subgroups came as a result of robustness testing,
which was a surprise to the researchers. Moreover, the researchers believe that lagged
variables played a major role in the performance of the model since they provided a historical
perspective of the financial situation that was essential for risk prediction.
Table 8. Performance of comparative model
Metric TabNet Logistic Regression Random Forest XGBoost Ensemble
Balanced Accuracy 0.782 0.694 0.743 0.764 0.796
Cohen's Kappa 0.693 0.567 0.632 0.663 0.702
F1 Score (Low Risk) 0.812 0.735 0.774 0.792 0.821
F1 Score (Medium Risk) 0.776 0.694 0.736 0.751 0.784
F1 Score (High Risk) 0.702 0.592 0.643 0.682 0.716
Average AUC-ROC 0.881 0.812 0.853 0.872 0.889
Ranking Loss Score 0.231 0.364 0.287 0.256 0.219

Source: Author’s Analysis results

4.5. Hypothesis Testing


The research results, indeed, very much support Hypothesis 1, revealing the very strong
predictive power of leverage and debt structure indicators. Hypothesis 2 gets part of the
support, where ROA_L1 takes the lead over lagged liquidity indicators in predictive
importance, though the time dynamics of risk signals are more intricate than hypothesised at
the start. Hypothesis 3 gets a very strong support, as TabNet is found to be a very significant
outperformer over classical methods in the area of non-linear relationships. Hypothesis 4 gets
a partial support, where STDTD indicates smaller firms as more significant than larger ones.
At last, Hypothesis 5 gets a support, with cash flow indicators offering predictive information
that is of value and goes beyond just the accounting measures.
4.6. Synthesis and Interpretation
Yet again, TabNet, an attention-based neural network, has proven efficient in predicting
financial distress for listed Vietnamese SMEs. TabNet performs better than classic models:
Logistic Regression finishes off with 69.4% balanced accuracy whereas TabNet delivers
78.2% balanced accuracy and is also outstanding with respect to F1-score in the high-risk
label identification. It is interpretable via sequential attention, offering transparent variable
importance evaluation, thereby allowing it to be readily adopted by practitioners and
regulatory bodies (Arthur, 2021; Nguyen et al., 2024).
The research identified the debt structure, more particularly the debt-to-equity (DE) ratio, as
the most important risk predictor. Due to the shadowing pattern caused by the nonlinear
threshold, the financial risk becomes noticeable at DE of 1.0 and further increase in value
12
from 2.0, thus supporting the theory of debt overhang and further underlining the sensitivity
of SMEs to leverage (Bui et al., 2021). The importance of the short-term debt to total debt
ratio (STDTD) associated with smaller enterprises confirms that, according to the corporate
life cycle theory, they are subject to rollover and interest rate risks (Ashraf et al., 2019). It is
interesting to note that the relatively decreased weight for the interest coverage ratio suggests
that it is imbalances in capital structure rather than servicing concerns over short term that
render the SMEs vulnerable (Mittal & Raman, 2020).
Following signalling theory, TabNet also finds accrual-based accounting ratios (e.g., ROA,
DE) to carry greater predictive power in emerging markets (Ashraf et al., 2019). On the other
hand, cash flow variables complement them when accounting figures underrate risk, thereby
supporting a multi-channel signalling perspective (Gryglewicz et al., 2022). The model
evaluates in steps, paralleling the theoretical evolution from profitability failure to liquidity
crises in distress (Chauhan, 2023; Matemane et al., 2024), thus furnishing an empirical basis
for the Vietnamese SME context.

5.

Tables and figures should be numbered (without headings) consecutively (in Arabic numbers,
Times New Roman 10, Bold, Centered, After 3 pt). The title (for both tables and figures)
should be placed above (see Figure 1, Table 1)
Table 1. Title

Table text – Times New Roman 9


Table borders – outside and titles – 1 ½ pt
Source: Times New Roman 9, Italic, Centered, Before 3.

Figures should be center aligned. Add figures as pictures with the highest possible resolution.
Use only black & white figures. All figures should be placed on portrait-oriented pages.
Figure 1. Title

Source: Times New Roman 9, Italic, Centered, Before 3.

13
Text2 (Times New Roman 10, Before 6 pt, After 6 pt, Justified)
Equations should be left-aligned in the first column of the single-row-two-columns table (see the model below).
The table with the equation should be set for 6 pt before and 6 pt after. The equations are consecutively
numbered, with Arabic numbers in parentheses, in the second column of the table, right-aligned (see the model
below).

(1)

References (Times New Roman 12, Bold, Before 12 pt, After 12 pt, Justified)
Text of references (Times New Roman 10, Regular, Justified, Hanging 1 cm)

INSTRUCTIONS FOR REFERENCES

Cited literature to include research papers from indexed journals in Scopus/ Web of science is highly
encouraged!

Citations of authors/works should be made in the text (NOT as footnotes) in the following general form: (Allen,
Wood, 2006, p. …). The page range is provided only in the citing inside the text. In the Reference list at
the end of the paper, only journal articles should have a page range.
The references and citations should follow the Harvard System Convention an APA Citation Style. All
references should be cited in the text.

For book: Author surname, Author initial. (Year). Title. Place: Publisher, ISBN.
Barker, R., Kirk, J. and Munday, R. J. (1988). Narrative Analysis. 3rd ed. Bloomington: Indiana University Press.

For book chapter: Chapter Author surname, Chapter Author initial. (Year). Chapter Title. – In: Book Author
surname, Book Author initial. (ed. if needed). Book Title. Place: Publisher.
Samson, C. (1970). Problems of Information Studies in History. – In: Stone, S. (ed.). Humanities Information
Research. Sheffield, CRUS.
2

14
For a journal article: Author surname, Author initial. (Year). Title. – Journal name, Vol. …, N …, p.p. …., ISSN.
Boughton, J. M. (2002). The Bretton Woods Proposal: an Indepth Look. – Political Science Quarterly, 42(6), p.
564-578.

For e-journals: Author surname, Author initial. (Year). Title. – Journal name, Vol. …., N …, p.p. …., available
at: <internet address>.
Kipper, D. (2008). Japan’s New Dawn. – Popular Science and Technology, [online] Available
at:<[Link]

OTHER SPECIFICS OF THE PAPER

Page Numbers (Times New Roman 8, Outside of margins)

Page Setup:
 Paper size – A4
 All pages – portrait orientation
 Margins – top 3 cm / bottom 3 cm / left 2.5 cm / right 2.5 cm
 Header – 3 cm
 Footer – 3 cm
 Different first page
 Different odd and even page

15

You might also like