CMA Inter | Module 9 | Data Processing, Organisation, Cleaning & Validation Paper 11 | Section B
CMA INTERMEDIATE | GROUP 2 | PAPER 11
Financial Management & Business Data Analytics
SECTION B – Business Data Analytics | 20% Weightage
MODULE 9
Data Processing, Organisation, Cleaning and Validation
Sections: 9.1 Development of Data Processing | 9.2 Functions of Data Processing | 9.3
Data Organisation & Distribution | 9.4 Data Cleaning & Validation
MODULE LEARNING OBJECTIVES
⬤ Understand the basic concepts of development of data processing.
⬤ Understand the basic concepts of functions of data processing.
⬤ Understand the basic concepts of data organisation and distribution.
⬤ Understand the basic concepts of data cleaning and validation.
9.1 Development of Data Processing
WHAT IS DATA PROCESSING (DP)?
Data processing (DP) is the process of organising, categorising and manipulating data in
order to extract information. Information in this context refers to valuable connections and
trends that address pressing issues. The capacity and effectiveness of DP have increased
manifold with technological development.
THREE PHASES IN THE HISTORY OF DATA PROCESSING
Phase Description
Manual DP Processing data without machine assistance. Small-scale operations only.
Still used today for data that is difficult to digitize (e.g., outdated texts or
documents).
Mechanical DP Uses mechanical tools (punch card machines), not modern computers.
Began in 1890 when the US Bureau of the Census used punch cards for
national census compilation. Faster and easier than manual processing.
Electronic DP Uses computers and cutting-edge electronics. Replaced both manual and
mechanical methods. Resulted in fewer errors and higher productivity.
Now widely used in industry, research and academia.
The Institute of Cost Accountants of India Page 1
CMA Inter | Module 9 | Data Processing, Organisation, Cleaning & Validation Paper 11 | Section B
DATA SCIENCE & FINANCE — 11 KEY AREAS
The relevance of data processing and data science in finance is increasing every day. The eleven
significant areas where data science plays an important role are:
Area Key Points
(i) Risk Analytics Identifies cybersecurity and financial risks. ML algorithms assess
customer reliability and lending risk. Creates dynamic real-time risk
models from transaction data.
(ii) Real-Time Enables instant response to consumer interactions. Dynamic data
Analytics pipelines and streams eliminate batch processing delays. Credit
ratings and transactions become more precise.
(iii) Customer Data Manages massive unstructured data using text analytics, data
Mgmt mining and NLP. Better understanding of market patterns and client
behaviour.
(iv) Consumer Machine learning enables customised service at scale. Insurance
Analytics firms use real-time analytics to calculate customer lifetime value
and cross-sell products.
(v) Customer Customers are grouped by geography, age, and buying patterns. ML
Segmentation algorithms assign relevance scores to attributes to assess customer
value.
(vi) Personalised NLP and voice recognition identify revenue opportunities and
Services enhance customer experience. Cross-selling facilitated by deep
understanding of customer needs.
(vii) Advanced Real-time evaluation of client interactions for effective
Customer Svc recommendations. Knowledge from each encounter improves future
interactions.
(viii) Predictive ML techniques learn from historical data to forecast future trends
Analytics (e.g., stock market). Deep learning performs better than shallow
learning as it doesn't require manual data preparation.
(ix) Fraud Detection AI and ML detect credit card fraud in real-time. Unusual patterns
(e.g., large purchase by frugal customer) trigger immediate card
suspension and notification.
(x) Anomaly Detects insider trading and account abnormalities using Recurrent
Detection Neural Networks and Long Short-Term Memory (LSTM) models.
Transformers used for next-generation solutions.
(xi) Algorithmic Unsupervised computers trade using algorithmic suggestions.
Trading Reinforcement Learning — model learns from penalties and
rewards. Eliminates human indecision and error. Operates in
fractions of a second.
The Institute of Cost Accountants of India Page 2
CMA Inter | Module 9 | Data Processing, Organisation, Cleaning & Validation Paper 11 | Section B
9.2 Functions of Data Processing
Data processing generally involves six key functions:
Function Description
(i) Validation Verifying whether data values come from an acceptable set (UNECE
2013). Ensures the final published data corresponds to quality
characteristics (Simon 2013).
(ii) Sorting Organising data into a meaningful order (ascending/descending) to aid
comprehension, analysis and visualisation. Can be applied to raw data or
aggregated information.
(iii) Aggregation Collecting and summarising data — individual rows are replaced with
summaries or totals. Data warehouses use aggregated data for fast
analytical queries.
(iv) Analysis Process of cleaning, converting and modelling data to obtain actionable
business intelligence. Extracts relevant information to support decision-
making.
(v) Reporting Gathering and structuring raw data into a consumable format to evaluate
organisational performance. Can be presented as text, graphs or charts.
(vi) Classifying data into categories to make it easier to use, safeguard and
Classification retrieve. Avoids duplications and reduces storage costs. Three methods:
content-based, context-based, user-based.
THREE METHODS OF DATA CLASSIFICATION (UNDER CLASSIFICATION
FUNCTION)
Method Description
Content-Based Examines and interprets files for sensitive data based on the actual
content within documents.
Context-Based Considers application, location, and creator as indirect markers of
sensitive information.
User-Based Relies on human selection and judgement by the end user during
document creation, editing, review or distribution.
The Institute of Cost Accountants of India Page 3
CMA Inter | Module 9 | Data Processing, Organisation, Cleaning & Validation Paper 11 | Section B
THREE RISK CATEGORIES IN DATA CLASSIFICATION
Risk Level Description & Examples
Low Risk Publicly accessible data. Simple recovery process. Examples: Public
website content, published annual reports.
Moderate Risk Non-public or internal data (within a business or its partners). Not too
mission-critical. Examples: Proprietary operating processes, cost of
products, corporate paperwork.
High Risk Sensitive or critical data for operational security. Difficult to retrieve if
lost. Examples: All secret, sensitive and essential data including financial
records.
STEPS FOR EFFECTIVE DATA CLASSIFICATION
✓ Understanding the current setup: Comprehensively review the location of current data and
applicable legislation before classifying.
✓ Creation of a data classification policy: Without an adequate policy, maintaining compliance is
practically impossible. This is the top priority.
✓ Prioritize and organize data: Once the policy is in place, categorize data based on sensitivity and
privacy using the optimal tagging method.
The Institute of Cost Accountants of India Page 4
CMA Inter | Module 9 | Data Processing, Organisation, Cleaning & Validation Paper 11 | Section B
9.3 Data Organisation and Distribution
DATA ORGANISATION
Definition
The classification of unstructured data into distinct groups. Raw data comprises variables'
observations.
Example: Arranging students' grades in different topics is a form of data organisation.
Key Points
As data volume grows, unstructured data becomes increasingly difficult to retrieve.
Techniques include: Classification, Frequency Distribution Tables, Image Representations,
Graphical Representations.
Allows arrangement of data in an easy-to-understand and manipulate manner.
IT workers use data organisation under the broader term 'data management'.
Structured data: Tabular, easily imported into databases.
Unstructured data: Raw and unformatted (e.g., text documents with scattered information).
Businesses implement data organisation to make better use of their data assets.
DATA DISTRIBUTION
Definition
A function that identifies and quantifies all potential values for a variable, as well as their
relative frequency (probability of how often they occur).
Any population with dispersed data is categorised as a distribution.
The primary benefit: estimation of the probability of any certain observation within a sample
space.
Statistics makes extensive use of data distributions. Raw data points are arranged into graphical
representations (histograms, box plots, pie charts) to give relevant information. Probability distribution
models are used to specify distinct sorts of random variables (discrete or continuous) in order to support
decision-making.
The Institute of Cost Accountants of India Page 5
CMA Inter | Module 9 | Data Processing, Organisation, Cleaning & Validation Paper 11 | Section B
TYPES OF DISTRIBUTION — SUMMARY TABLE
Distribution Description Example
DISCRETE
DISTRIBUTIONS
Binomial Quantifies the chance of obtaining a Eg:
specific number of successes or failures. Coin toss —
Applies to two mutually exclusive and P(Head) = 1/2,
exhaustive classes (success/failure, P(Tail) = 1/2.
accept/reject).
Poisson Discrete probability distribution for the Eg:
number of events in a given time period. Number of
Applies to attributes that can take huge flaws,
values but in practice take tiny ones. mistakes,
accidents,
absentees.
Hypergeometric Assesses probability of successes in (n) Eg:
trials without replacement from a large Quality
population (N). Similar to binomial but sampling from
probability of success is constant. a batch without
replacement.
Geometric Assesses the probability of the first Eg:
success occurring. Extension: negative Choosing
binomial distribution. hockey players
randomly until
an Olympic
participant is
found.
CONTINUOUS
DISTRIBUTIONS
Normal Bell-shaped (Gaussian) curve with Eg:
highest frequency at centre. Values on Heights,
either side of the target value have equal weights, exam
likelihood. scores.
Lognormal Variable x follows lognormal Eg:
distribution if its natural logarithm ln(x) Stock price
is normally distributed. Sum of random returns,
variables approaches normal as sample income
size rises. distribution.
F Distribution Used to examine equality of variances Eg:
between two normal populations. ANOVA tests,
Asymmetric with minimum value of 0 comparing
and no maximum value. group
variances.
The Institute of Cost Accountants of India Page 6
CMA Inter | Module 9 | Data Processing, Organisation, Cleaning & Validation Paper 11 | Section B
Distribution Description Example
Chi-Square Occurs when squared independent Eg:
standard normal variables are added y = Z1² + Z2²
together. Symmetrical, constrained + ... + Zn²
below zero. Approaches normal
distribution as degrees of freedom
increase.
Exponential Most commonly used continuous Eg:
distribution. Represents products with a Time between
constant failure rate. Closely connected failures in
to Poisson distribution. equipment.
T (Student's) Bell-shaped, symmetrical about its mean. Eg:
Used for hypothesis testing and Testing sample
confidence intervals when standard means when
deviation is unknown. population SD
is unknown.
The Institute of Cost Accountants of India Page 7
CMA Inter | Module 9 | Data Processing, Organisation, Cleaning & Validation Paper 11 | Section B
9.4 Data Cleaning and Validation
DATA CLEANING
DEFINITION
Data cleaning is the process of correcting or deleting inaccurate, corrupted, improperly
formatted, duplicate or insufficient data from a dataset. Incorrect data renders outcomes and
algorithms untrustworthy despite their apparent accuracy.
Data Cleaning Data Transformation
Process of REMOVING irrelevant data from Process of CHANGING data from one
a dataset. format/structure to another. Also called data
wrangling or data munging.
5 STEPS OF DATA CLEANING
Step Action Key Points
Step Remove Duplicates ● Eliminate duplicate observations (common when merging
1 & Irrelevant Data data from multiple sources).
● Remove irrelevant observations not related to the study
objective.
● De-duplication is one of the most important
considerations in this step.
● Reduces distractions and makes the dataset more
manageable.
Step Fix Structural ● Detect unusual naming standards, typos or wrong
2 Errors capitalisation during data transfer.
● Merge inconsistently labelled entries: e.g., 'N/A' and 'Not
Applicable' should be unified.
● Prevent mislabelled classes or groups.
Step Filter Unwanted ● Identify observations that do not fit the data at first glance.
3 Outliers ● Remove outliers caused by erroneous data input — but
remember: an outlier is not automatically an error.
● Consider deleting an outlier only if it appears unrelated to
the analysis.
Step Handle Missing ● Missing data cannot be ignored — many algorithms reject
4 Data missing values.
● Option 1: Drop observations with missing values (risk:
loss of information).
The Institute of Cost Accountants of India Page 8
CMA Inter | Module 9 | Data Processing, Organisation, Cleaning & Validation Paper 11 | Section B
Step Action Key Points
● Option 2: Impute missing values based on other
observations (risk: assumptions may compromise
integrity).
Step Validation & QA ● Does the data make sense?
5 ● Does the data adhere to regulations applicable to its field?
● Does it verify or contradict your working hypothesis?
● Can data patterns assist in formulating the next theory?
● If not, is this due to a data quality issue?
BENEFITS OF DATA CLEANING
✓ Error correction when numerous data sources are involved.
✓ Fewer mistakes result in happier customers and less irritated employees.
✓ Capability to map the many functions and planned uses of your data.
✓ Monitoring mistakes and improving reporting makes it easier to repair inaccurate data in future
applications.
✓ Using data cleaning technologies results in more effective corporate procedures and speedier
decision-making.
FOUR CHARACTERISTICS OF QUALITY DATA
Validity Accuracy Completeness Consistency
Ensures data Data is correct and All necessary data is Data is uniform
conforms to defined free from errors. present with no across different
formats and rules. missing values. datasets and systems.
DATA VALIDATION
DEFINITION
Data validation is a crucial component of any data management process. It ensures that the
initial data is valid so that outcomes are accurate. Using validation criteria to purify data
before usage mitigates 'garbage in, garbage out' problems.
6 TYPES OF DATA VALIDATION
Validation Type Description Example
Data Type Verifies that entered data has the appropriate data A field that only
Check type. Rejects characters that do not match the accepts numbers
expected type (e.g., letters in a numeric field). rejects alphabets.
The Institute of Cost Accountants of India Page 9
CMA Inter | Module 9 | Data Processing, Organisation, Cleaning & Validation Paper 11 | Section B
Validation Type Description Example
Code Check Verifies that a field's value is picked from a Postal code
legitimate set of options or follows specific compared to a list of
formatting rules. valid codes; NIC
industry codes.
Range Check Determines whether input data falls inside a Latitude: -90 to 90;
specified range. Values outside the range are Longitude: -180 to
invalid. 180.
Format Check Ensures data adheres to a defined format for Date format: 'YYYY-
consistency. MM-DD' or 'DD-
MM-YYYY'.
Consistency Logical check that verifies data is entered in a Delivery date must
Check consistent manner across related fields. be after shipment
date.
Uniqueness Guarantees that unique items (like PAN or email PAN number, Email
Check IDs) are not entered into a database multiple ID — must be
times. unique in the
database.
SOLVED CASE STUDY — MAITREYEE (DATA CLEANING)
Case Facts
Maitreyee is a data analyst with a financial organisation. She has received a large amount of
data and plans to use statistical techniques, but finds the data is not cleaned. Before applying
any data analysis tools, cleaning is essential.
Steps Maitreyee should follow:
1. Step 1: Remove duplicates and irrelevant data
2. Step 2: Fix structural errors
3. Step 3: Filter unwanted outliers
4. Step 4: Handle missing data
5. Step 5: Validation and QA — Does data make sense? Does it adhere to regulations?
Does it verify the hypothesis?
Benefits of Clean Data:
Validity | Accuracy | Completeness | Consistency
The Institute of Cost Accountants of India Page 10
CMA Inter | Module 9 | Data Processing, Organisation, Cleaning & Validation Paper 11 | Section B
EXERCISE — Exam Practice Questions
A. Multiple Choice Questions (MCQs)
1. Data science plays an important role in:
(a) Risk analytics
(b) Customer data management
(c) Consumer analytics
(d) All of the above ✓
2. The primary benefit of data distribution is:
(a) The estimation of the probability of any certain observation within a sample
space ✓
(b) The estimation of the probability within a non-sample space
(c) The estimation of the probability within a population
(d) None of the above
3. Binomial distribution applies to attributes:
(a) That are categorised into two mutually exclusive and exhaustive classes ✓
(b) That are categorised into three mutually exclusive classes
(c) That are categorised into less than two classes
(d) That are categorised into four mutually exclusive classes
4. The geometric distribution assesses:
(a) The probability of the occurrence of the first success ✓
(b) The probability of the second success
(c) The probability of the third success
(d) None of the above
5. The probability density function describes:
(a) The characteristics of a random variable ✓
(b) The characteristics of a non-random variable
(c) The characteristics of a random constant
(d) The characteristics of a non-random constant
6. Which phase of data processing used punch card machines?
(a) Manual DP
The Institute of Cost Accountants of India Page 11
CMA Inter | Module 9 | Data Processing, Organisation, Cleaning & Validation Paper 11 | Section B
(b) Mechanical DP ✓
(c) Electronic DP
(d) None of the above
7. Data cleaning is different from data transformation because:
(a) Data cleaning removes irrelevant data; transformation changes data format or
structure ✓
(b) Data cleaning is more expensive
(c) Transformation removes duplicate data
(d) None of the above
8. Which validation type ensures that a date field follows 'YYYY-MM-DD' format?
(a) Range check
(b) Code check
(c) Format check ✓
(d) Uniqueness check
B. True or False
No. Statement Answer
1. Data validation is a process which ensures the correspondence of the final TRUE
data with a number of quality characteristics.
2. Data analysis is described as the process of cleaning, converting and TRUE
modelling data to obtain actionable business intelligence.
3. Financial data such as revenues, accounts receivable and net profits are often TRUE
summarised in a company's data reporting.
4. Structured data consists of tabular information that may be readily imported TRUE
into a database.
5. Data distribution is a function that identifies and quantifies all potential TRUE
values for a variable and their relative frequency.
6. Manual DP is no longer in use today. FALSE
7. The Poisson distribution is a continuous probability distribution. FALSE
8. Chi-square distribution is constrained below zero and approaches normal TRUE
distribution as degrees of freedom increase.
C. Fill in the Blanks
The Institute of Cost Accountants of India Page 12
CMA Inter | Module 9 | Data Processing, Organisation, Cleaning & Validation Paper 11 | Section B
No. Statement Answer
1. Data may be classified as Restricted, ________ or Public by Private
an entity.
2. Data organisation is the ________ of unstructured data into Classification
distinct groups.
3. Classification, frequency distribution tables, ________, Image representations
graphical representations etc. are examples of data
organisation techniques.
4. The t distribution is a probability distribution with a bell Mean
shape that is symmetrical about its ________.
5. ________ is the process of correcting or deleting inaccurate, Data cleaning
corrupted, improperly formatted, duplicate or insufficient
data from a dataset.
6. The history of data processing can be divided into ________ Three
phases.
7. ________ trading happens when an unsupervised computer Algorithmic
uses an algorithm to suggest trades on the stock market.
8. The ________ distribution quantifies the chance of a certain Poisson
number of events occurring in a given time period.
D. Short Essay Questions
1. Briefly discuss the role of data analysis in fraud detection.
2. Discuss the difference between discrete distribution and continuous distribution.
3. Write a short note on binomial distribution.
4. What is the significance of data cleaning?
5. Write a short note on 'predictive analytics'.
6. What is data organisation? Give examples of data organisation techniques.
7. Distinguish between data cleaning and data transformation.
E. Essay Questions
1. Elaborately discuss the functions of data processing.
2. Elaborately discuss the various steps involved in data cleaning.
3. Discuss the benefits of data cleaning.
The Institute of Cost Accountants of India Page 13
CMA Inter | Module 9 | Data Processing, Organisation, Cleaning & Validation Paper 11 | Section B
4. How are data processing and data science relevant for finance? Discuss the 11 key areas.
5. Discuss the steps for effective data classification.
6. Discuss the types of data distribution with examples.
7. Explain the six types of data validation with suitable examples.
F. Unsolved Case Study
Case Study — Arjun & Akansha Limited
Arjun is a data analyst working with Akansha Limited, a company that retails FMCG
products through both online and offline channels. Over the years, the company has
accumulated a huge amount of data. The management is puzzled about how to bring this
data into a usable format. Arjun is entrusted with this responsibility.
Required: Make suggestions on how Arjun can bring the data into usable format.
Discuss relevant steps and techniques for data organisation, cleaning and validation.
The Institute of Cost Accountants of India Page 14