Accounting Info & Stock Returns in Vietnam
Accounting Info & Stock Returns in Vietnam
This paper studies the relationship between accounting information reflected in financial state-
ments and stock return in Vietnam Stock Market. The authors propose a research model to define
the relationship. Besides, machine learning algorithms are used in researching and forecasting,
with data obtained from observed firms during the period from 2009 to 2020. Research results
show that gradient boosting algorithm has the best self-reporting performance and finan-
cial ratios also have great impact on stock returns, including operating income growth, stock
earnings volatility, dividend yield, earnings before tax-to-equity ratio, cash holding ratio, and
accrual quality. Based on the research results, the authors make some recommendations for
investors, firms, and policy makers.
Keywords: stock return, machine learning, decision tree, gradient boosting, elastic net, accoun-
ting information, accrual quality
Este artículo estudia la relación entre la información contable reflejada en los estados financie-
ros y el rendimiento de las acciones en el mercado de valores de Vietnam. Se propone un modelo
de investigación para definir la relación. Además, los algoritmos de aprendizaje automático se
utilizan en la investigación y la previsión, con datos obtenidos de empresas observadas durante
[Link]
Contabilidad y Negocios (17) 33, 2022, pp. 94-118 / e-ISSN 2221-724X
Accounting information and stock returns in Vietnam securities market 95
el período de 2009 a 2020. Los resultados de la investigación muestran que el algoritmo gradient
boosting tiene el mejor rendimiento de autoinforme y los índices financieros también tienen un
gran impacto en la rentabilidad de las acciones, al incluir el crecimiento de los ingresos operati-
vos, la volatilidad de las ganancias de las acciones, el rendimiento de los dividendos, la relación
entre las ganancias antes de impuestos y el capital, la relación de tenencia de efectivo, y la
calidad de la acumulación. Sobre la base de los resultados de la investigación, se hacen algunas
recomendaciones para inversionistas, empresas y formuladores de políticas.
Palabras clave: rentabilidad de acciones, aprendizaje automático decision tree, gradient boos-
ting, elastic net, información de cuenta, calidad devengada
Este artigo estuda a relação entre as informações contábeis refletidas nas demonstrações
financeiras e o retorno das ações no mercado de ações do Vietnã. Os autores propõem um
modelo de pesquisa para definir a relação. Além disso, algoritmos de aprendizado de máquina
são usados em pesquisas e previsões, com dados obtidos de empresas observadas durante o
período de 2009 a 2020. Os resultados da pesquisa mostram que o algoritmo gradient boosting
tem o melhor desempenho de autorrelato e os índices financeiros também têm grande impacto
nos retornos das ações, incluindo crescimento da receita operacional, volatilidade do lucro
das ações, rendimento de dividendos, lucro antes do índice de impostos sobre o patrimônio
líquido, índice de retenção de caixa e qualidade de acumulação. Com base nos resultados da
pesquisa, os autores fazem algumas recomendações para investidores, empresas e formula-
dores de políticas.
1. Introduction
Stock returns have been a topic of scientific research interest for many years (Ball
& Brown, 1968; Basu, 1983; Freeman, 1987; Collins & Kothari, 1989; Easton & Harris,
1991) have built yield models from income and return variables, by testing the rela-
tionship between current accounting earning divided by stock price at the beginning
of the period and stock return. Investors aim to maximize the value of their assets
by investing in firm shares as they offer the most significant potential for long-term
returns based on a trade-off of return and risk. Anwaar (2016) believes that a firm’s
financial and accounting information can be used effectively to evaluate investment
opportunities for investors. A firm’s information is divided into internal information
and external information, based on the source of the information (Emamgholipour et
al., 2013). As the name implies, external information is taken from the stock market,
while internal information appears on financial statements. Therefore, investors can
rely on accounting information on financial statements to evaluate stock returns over
time when making investment decisions.
Recently, there is a promising research field which focuses more on artificial intel-
ligence techniques and can handle non-linearity and stochasticity. Improvements in
technology and computer science have spread implementations to a wider audience,
spurring the growth of research in machine learning. Financial market prediction
research is done by applying several machine learning algorithms, including gradient
boosting, support vector machine, and random forest to predict price, profitability,
direction and volatility of securities, stocks and commodities indicators (Henrique
et al., 2019).
to determine which indicators are suitable for making predictions. Therefore, there
is a need to study the impact of accounting information on stock returns on Vietnam
Stock Market.
The study aims to give an overview of the relationship between accounting infor-
mation, through relevant financial ratios, and stock returns prediction by answering
the research question and the following questions: “How is the accounting information
on financial statements reflected in financial ratios associated with the stock returns
on Vietnam Stock Market?”. Furthermore, the following detailed questions are formu-
lated to answer whether certain financial ratios have a stronger association with stock
returns than others. To determine the association, we use and compare machine learn-
ing models to each other to answer the questions: Which financial ratios have the
strongest association with stock returns? Which applied machine learning model has
the highest predicting accuracy?
Literature review
Using accounting information is perhaps the most common way to evaluate firm
performance and provides essential inside information about firm value and market
capitalization. Accounting information is widely used for many years all over the world,
especially to predict the future performance of firms. Additionally, it can be used as
a tool to predict stock returns in the stock market (Lewellen, 2004). It is also used to
determine a firm’s creditworthiness.
Financial theory suggests that financial ratios can provide an insight into a firm’s
operating and financial performance by rearranging accounting information on finan-
cial statements such as returns and firm value, its ability to allocate cash flows to
investments, its liquidity and its debt size. Financial ratios are also used for future
performance prediction. For example, they are used as explanatory variables in predic-
tive modeling to predict financial distress, failure, and bankruptcy (Sun et al., 2011;
Zięba et al., 2016; Kim & Upneja, 2014). Similarly, financial ratios are used in modeling
the relationship between them and stock returns, by using multiple or sets of indi-
vidual financial variables or using statistical methods or machine learning techniques
(Emamgholipour et al., 2013; Lewellen, 2004). Although most of the previous studies are
successful in modeling the relationship between financial ratios and firm returns, they
lack determination about the importance features of these ratios used to evaluate firm
performance.
In Vietnam, there have been a number of studies by Hai et al. (2015), Ha et al. (2018),
Hung et al. (2018), and Hung and Van (2020) showing the impact of accounting infor-
mation on stock returns. However, studies which are based on machine learning algo-
rithms and the importance of financial indicators on stock returns have not been yet
considered. Results from previous studies show evidence of relationship between
accounting information and returns. Yet, those studies are based on various theories
and models. Each model has its own advantages and disadvantages, and the data
collection methods are different. This study will determine which algorithm is most
suitable in determining the impact of accounting information on stock returns, and
which financial ratios have the strongest impact.
3. Research methodology
3.1. Research model
From the studies presented in the literature review such as Lewellen (2004), Delen et
al. (2013), Emamgholipour et al. (2013), Musallam (2018), Hai et al. (2015), Ha et al. (2018),
Hung et al. (2018), Hung and Van (2020), and Metsomäki (2020), the authors establish
the model as follows:
Stock return is a ratio that measures the percentage of a stock’s net income per
dollar of invested capital. The formula for the stock returns is the appreciation in the
price plus any dividends paid, divided by price at the beginning of the period. This
calculation method is frequently mentioned in studies on stock returns and in the
original model of Easton and Harris (1991), which is:
Pit-1 , Pit : price of share at the end of year t and t-1, respectively
Dit : dividend per share for the year t.
Linear regression: As the first linear model, linear regression with the ordinary least
square method is usually implemented. its aim is to minimize the sum of squares
between the true and estimated values by fitting the linear model with coefficients
(Pedregosa et al., 2011).
Ridge regression and Lasso regression are two regression models that apply regu-
larization techniques to avoid overfitting. Overfitting is a phenomenon where the
model only fits well on training dataset but does not predict well on test data. This is
often the case when training machine learning models. This phenomenon has a nega-
tive effect and leads to inapplicable models because the predictions turn wrong when
applied in practice. There are many causes of overfitting. One of the common causes is
that the training dataset and the test dataset have different distribution, making the
rules learned in the training dataset no longer valid in the test dataset, or it is because
when a model has too many parameters, its data is not representative. Regularization
is a technique to avoid overfiting by adding a regularization term to the loss function.
Usually, this term is in the form of L1 norm or L2 norm of the coefficients. They are
called ridge regression and Lasso regression respectively.
For these regressions, we need to fit coefficient α to find a coefficient that is best
for each dataset. If data is severely overfitting, it is necessary to reduce overfitting by
increasing the effect of the regularization term by increasing coefficient α. If the model
is not overfitting, α can be close to 0. If α = 0, the regression equation is equivalent to
multivariable linear regression.
Elastic net: To add more explanatory power to the family of linear models, elastic
net (Zou & Hastie, 2005) was applied. Firstly, it is a continuation of linear regression
models trained with Lasso’s L1 penalty and Ridge’s L2 penalty. Combining the penalties
of both methods in one model produces a regular, competitive model in which the
weight of parameters is non-zero (Pedregosa et al., 2011).
and attributes can be of different data types such as binary, nominal, ordinal, quanti-
tative data. To determine which variable is classified first, which variable is classified
later, the weight of evidence (Entropy) for each variable is calculated, the higher the
information value, the more classification information the variable carries.
The other two applied models are grouped together, although they are not as similar
as previous models, they use the same method to evaluate the relationship between
financial ratios and stock returns.
Support vector machine (SVM): SVM is a binary classification algorithm, SVM takes input
data and classifies them into two different classes. Given a set of training examples of two
given categories, SVM algorithm builds a SVM model to classify other examples into those
two categories. SVM builds a hyperplane to classify a dataset into two separate classes.
To do this, first SVM will construct a hyperplane or a set of hyperplanes in a multi-dimen-
sional or infinite-dimensional space, which can also be used for classification, regres-
sion, or other tasks. For the best classification, it is necessary that optimal hyperplane is
located as far away from the data points of all classes (margins) as possible, because the
larger the margins, the smaller the generalization errors of the classification algorithm.
There are three types of importance scores for more advanced features that are
also implemented from model coefficients as part of the linear model, from decision
model and permutation importance, which are described in detail by Pedregosa et al.
(2011). We describe three important features as follows:
Decision tree feature importance: Decision tree algorithms, like the CART algorithm
used in this study, are provided in scikit-learning implementations with feature impor-
tance scores, based on criterion reduction used to select split points. This approach
is applied to decision tree model and all the tree-based combination methods such as
random forest, gradient boosting, and AdaBoost.
1
RMSEi
Wi = , for i = 1,...n
1
∑ j RMSE
n
Evaluating a regression model is perhaps the most important part of building a super-
vised machine learning model. This determines whether a model is ready for deploy-
ment. Indicators only show numbers, but if examined carefully, they can tell us about
feature selection and feature engineering. In this study, we will look at the regression
metrics and discover why we use these following regression metrics: mean absolute
error, mean squared error and root mean squared error.
Mean absolute error (MAE) is a measure of the mean absolute error between
predicted and actual values.
1 n
MAE = ∑ yi - yiµ
n i=1
Basically, MAE is a L1 norm. The smaller MAE, the smaller the distance between
predicted and actual values, and the better the model. However, MAE value does not
cover unit difference.
The mean squared error (MSE) of an estimator measures the average squares of
errors – that is, the average squared difference between predicted values and actual
values. MSE is a risk function, corresponding to the expected value of squared error
loss or quadratic loss. MSE is the second moment (about the origin) of errors.
1 n
( )
2
µ
MAE =
n
∑ i=1
yi - yi
Root mean squared error (RMSE) is the standard deviation of residuals (prediction
errors). Residuals are a measure of how far from the regression line data points are. RMSE
is a measure of how spread-out residuals are. In other words, it tells you how concentrated
data is around the line of best fit. RMSE is commonly used in forecasting and regression
analysis to verify experimental results and is also a measure of how effective models are,
by measuring the difference between predicted and actual values. The smaller RMSE, the
smaller the errors, the higher the level of reliability estimation of models.
1 n
( )
2
µ
RMSE =
n
∑ i=1
yi - yi
This study uses data collected from Vietnam Stock Exchange over the period of
2009-2020. Data is collected from audited financial statements of listed firms after
excluding firms in banking, securities and insurance sectors. After determining the
indicators, there are 3397 observations to be analyzed and forecasted, presented in
table 1 by year and by industry.
Based on figure 1, the stock return of firms witnessed huge fluctuations over the
period of 2009-2010, especially in 2010, stock return decreased to -9%, while during
2011-2013, it increased to 81% in 2013. Then from 2014 to 2017, the return was stable
at around 20%. In 2019, the stock return dropped to -15%, but by 2020, it increased
dramatically to 66%.
120% 120%
100%
80% 81%
66%
60%
40%
32%
28% 27%
20% 22% 21%
0% 4% 2%
2009 -9%2011 2013 2015 2017 2019-15% 2021
-20%
-40%
Figure 2 shows observed stock returns by industry in the period of 2009-2020, in which
the industry with the highest stock return is the health industry, followed by the consumer
goods industry. The industries with low stock return are technology and industrial.
Statistical data in table 2 shows that the average stock return is 27.0%, the standard
deviation is 54.4%, the lowest return is -59.9% and the highest is 244.9%. The account-
ing information is reflected through 11 financial ratios in terms of mean, standard devi-
ation, minimum value and maximum value.
With MAE measurement, AdaBoost algorithm has the highest value of 0.498, which
proves that this algorithm has the lowest measurement performance, while SVM algo-
rithm has a value of 0.339, which proves that it has the best measuremance perfor-
mance among these ten algorithms in applied dataset. Meanwhile, with MSE and RMSE
measurement, the algorithm with the lowest value, 0.208 and 0.456 respectively, is
gradient boosting algorithm. Thus, gradient boosting algorithm has the best predic-
tion performance among the algorithms used in predicting the impact of accounting
information from financial indicators on stock returns.
A2 0.093
A5 0.063
A8 0.053
A7 0.040
A6 0.022
A9 0.016
A4 0.015
A11 0.012
A3 0.011
A10 0.010
A1 0.003
Next, we used decision tree algorithm group including random forest, decision tree,
AdaBoost, gradient boosting. Feature importance values of financial ratios are shown
in table 5 and their permutation feature importance values are presented in figure 5.
A10 0.201
A5 0.198
A6 0.098
A2- 0.095
A8 0.081
A7 0.064
A3 0.061
A11 0.056
A1 0.054
A9 0.048
A4- 0.045
As shown in figure 5, the three financial indicators that have the highest feature
importance score are accrual quality, operating income growth and stock earnings
volatility, with a value of 0.201, 0.198 and 0.098, respectively. The three financial ratios
with the lowest importance score are fixed assets turnover (0.045), earnings manage-
ment (0.0048) and debt-to-equity ratio (0.054).
The feature importance of support vector machine model and K-nearest neighbors
model are calculated as permutation feature importance. The initial values provided by
this technique are shown in table 6. The result is the average weighted importance score
for each financial indicator as determined in column F (Fused). The value of permutation
feature importance is much smaller than feature importance and coefficients as feature
importance presented in the previous sections, therefore they cannot be directly compared
with each other. However, importance feature rankings from highest to lowest can still
be presented as the average weighted importance score for each financial indicator.
0.003
0.002
0.001
0.000
A4 A9 A10 A11 A1 A6 A3 A5 A2 A8 A7
effect on stock returns, and higher net profit margin generates higher profits for the
next period. Previous research has also found that dividend yield can predict stock
returns (Fama & French, 2021; Lewellen, 2004). Similarly, the research results of Musal-
lam (2018) show that dividend yield is a highly important financial ratio, with positive
coefficient values, they show a positive association with stock returns.
12
10
0
A1 A2 A3 A4 A5 A6 A7 A8 A9 A10 A11
Research results of Musallam (2018), and Emamgholipour et al. (2013) indicate that
earnings per share has a positive and significant association with stock returns. This
result is consistent with these following research: Easton and Harris (1991), Ball and
Brown (1968), Basu (1983), and Hung et al. (2018). Stock earnings volatility also has
an important impact on stock returns, this result is consistent with that of Freeman
(1987), Collins and Kothari (1989), and Easton and Harris (1991). Besides, consistent with
a research by Hung and Van (2020), accrual quality also has a significant impact on
stock returns.
However, there are other factors that can affect stock returns and announced
accounting information has not yet strongly affected stock return volatility
on Vietnam Stock Market. Before making investment decisions, investors can
consider earnings before tax-to-equity ratio, dividend yield because these
indicators have an important and positive impact on stock returns. Investors
should also refer to other additional information to make the best decisions.
Empirical results of other indicators like cash holding ratio, on the relationship
between earnings and stock returns, have suggested the difference between
firms with the same earnings rate.
Contribución de autores
References
Altman, N. S. (1992). An introduction to kernel and nearest-neighbor nonparametric
regression. The American Statistician, 46(3), 175-185. [Link]
031305.1992.10475879
Anwaar, M. (2016). Impact of firms performance on stock returns (Evidence from listed
companies of ftse-100 index London, UK). Global Journal of Management and
Business Research, 16(1), 1-10.
Ball, R., & Brown, P. (1968). An empirical evaluation of accounting income numbers.
Journal of Accounting Research, 6(2), 159-178. [Link]
Barnes, P. (1987). The analysis and use of financial ratios. Journal of Business Finance
dan Accounting, 14(4), 449-461. [Link]
Basu, S. (1983). The relationship between earnings’ yield, market value and return for
NYSE common stocks: Further evidence. Journal of financial economics, 12(1),
129-156. [Link]
Belson, W. A. (1959). Matching and prediction on the principle of biological classifica-
tion. Journal of the Royal Statistical Society: Series C (Applied Statistics), 8(2),
65-75. [Link]
Collins, D. W., & Kothari, S. (1989). An analysis of intertemporal and cross-sectional
determinants of earnings response coefficients. Journal of Accounting and
Economics, 11(2-3), 143-181. [Link]
Dechow, P. M., & Dichev, I. D. (2002). The quality of accruals and earnings: The role
of accrual estimation errors. The Accounting Review, 77, 35-59. [Link]
org/10.2308/accr.2002.77.s-1.35
Delen, D., Kuzey, C., & Uyar, A. (2013). Measuring firm performance using financial ratios:
A decision tree approach. Expert Systems with Applications, 40(10), 3970-3983.
[Link]
Easton, P. D., & Harris, T. S. (1991). Earnings as an explanatory variable for returns. Jour-
nal of Accounting Research, 29(1), 19-36. [Link]
Emamgholipour, M., Pouraghajan, A., Tabari, N. A. Y., Haghparast, M., & Shirsavar, A.
A. A. (2013). The effects of performance evaluation market ratios on the stock
return: Evidence from the Tehran stock exchange. International Research Journal
of Applied and Basic Sciences, 4(3), 696-703. [Link]
Fama, E. F., & French, K. R. (2021). Dividend yields and expected stock returns. Univer-
sity of Chicago Press.
Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Dubourg, V.
(2011). Scikit-learn: Machine learning in Python. The Journal of Machine Learning
Research, 12, 2825-2830.
Quinlan, J. R. (1986). Induction of decision trees. Machine Learning, 1(1), 81-106. https://
[Link]/10.1007/BF00116251
Quinlan, J. R. (1996). Bagging, boosting, and C4. 5. En American Association for Artificial
Intelligence (Ed.), AAAI’96: Proceedings of the thirteenth national conference on
artificial intelligence. Vol 1 (pp. 725-730). AAAI Press.
Sun, J., Jia, M.-y., & Li, H. (2011). AdaBoost ensemble for financial distress predic-
tion: An empirical comparison with data from Chinese listed companies.
Expert Systems with Applications, 38(8), 9305-9312. [Link]
eswa.2011.01.042
Zięba, M., Tomczak, S. K., & Tomczak, J. M. (2016). Ensemble boosted trees with synthe-
tic features generation in application to bankruptcy prediction. Expert Systems
with Applications, 58, 93-101. [Link]
Zou, H., & Hastie, T. (2005). Regularization and variable selection via the elastic net.
Journal of the Royal Statistical Society: Series B (Statistical Methodology), 67(2),
301-320. [Link]