Part I - Research Methodology - (Common for all Candidate)
Meaning of Research: Objectives of Research, Types of Research, Research Process,
Problem Statement, Research Design, Approaches to Research - Quantitative,
Qualitative Approach, Exploratory, Confirmatory Research, Experimental and
Theoretical Research.
Problem Formulation: Conducting Literature Review, Information Sources (Books,
monographs, reviews, blogs, etc.), Information Retrieval, Role of libraries in
Information Retrieval, Tools for identifying literature (digital resources and print media),
Indexing and abstracting services, Citation indexes, Summarizing the Review, Critical
Review, Identifying Research Gap, Conceptualizing and Hypothesizing the research
gap.
Research Design: Experimental / Simulation/ Theoretical /Empirical Research, Cause
effect relationship, Development of Hypothesis, Measurement Systems Analysis, Validity
and Reliability, Statistical Design of Experiments, Field Experiments, Data/Variable
Types & Classification, Data collection - Methods and Tools.
Data Analysis and Interpretation: Sampling, Sampling Error, Statistical
Methods/Tools - Measures of Central Tendency and Variation, Test of Hypothesis -
z test, t test, F test, ANOVA, Chi square, correlation and regression analysis, Error
Estimation.
Writing Research Articles and Thesis: Data Presentation - Types of tables and
illustrations, Guidelines for writing the abstract, introduction, methodology, results
and discussion, conclusion sections of a manuscript. References - Styles and
methods, Citation and listing system of documents. Plagiarism. Ethical considerations in
Research
Meaning of Research: Objectives of Research, Types of Research, Research Process,
Problem Statement, Research Design, Approaches to Research - Quantitative,
Qualitative Approach, Exploratory, Confirmatory Research, Experimental and
Theoretical Research.
1. Meaning of Research
Definition:
Research is a systematic and scientific process of collecting, analyzing, and interpreting
data to answer questions or solve problems.
Purpose:
To discover new knowledge, verify existing knowledge, and develop new theories,
models, or methods.
Key Features:
o Systematic and organized
o Based on empirical evidence
o Uses scientific methods
o Objective and logical
o Aims at discovery and interpretation of facts
2. Objectives of Research
To gain familiarity with a phenomenon (exploratory studies)
To describe characteristics or functions of a situation (descriptive research)
To determine cause-and-effect relationships (analytical or experimental research)
To test hypotheses or theories
To develop new tools, methods, or theories for better understanding
To solve practical problems and improve existing processes or systems
3. Types of Research
Research can be classified based on purpose, approach, and nature of data:
a) Based on Purpose:
Basic (Pure) Research:
Aims to increase scientific knowledge without direct practical application (e.g.,
developing a new theory in physics).
Applied Research:
Focused on solving real-world problems (e.g., developing a new drug).
b) Based on Nature of Data:
Quantitative Research:
Uses numerical data and statistical analysis (e.g., survey, experiments).
Qualitative Research:
Uses non-numerical data to understand concepts, experiences, or meanings (e.g.,
interviews, case studies).
c) Based on Method:
Descriptive Research:
Describes existing phenomena without manipulation.
Analytical Research:
Uses existing data to make critical evaluations.
Experimental Research:
Involves manipulation of variables under controlled conditions.
Exploratory Research:
Conducted to explore a new area where little is known.
Historical Research:
Studies past events using records and documents.
4. Research Process
A typical research process follows these steps:
1. Identification of Research Problem
o Selecting an area of interest or gap in knowledge.
2. Review of Literature
o Studying existing research to identify what has already been done.
3. Formulation of Research Problem and Hypothesis
o Defining the research question and expected outcomes.
4. Research Design
o Planning methods, tools, and techniques for data collection and analysis.
5. Data Collection
o Gathering information through experiments, surveys, interviews, etc.
6. Data Analysis and Interpretation
o Using statistical or qualitative methods to draw conclusions.
7. Report Writing and Presentation
o Presenting findings in a structured report or thesis.
5. Problem Statement
Definition:
A clear and concise description of the issue or gap that the research aims to address.
Importance:
It guides the direction of the study and defines the objectives and scope.
Qualities of a Good Problem Statement:
o Clear and specific
o Researchable
o Feasible in terms of time and resources
o Significant for knowledge advancement
6. Research Design
Definition:
A blueprint or framework that outlines how research will be conducted.
Purpose:
Ensures that evidence obtained enables the researcher to answer the research question as
clearly as possible.
Components:
o Selection of research methods
o Sampling techniques
o Data collection instruments
o Plan for data analysis
Types of Research Design:
o Exploratory Design: To explore a new area or identify problems.
o Descriptive Design: To describe phenomena or characteristics.
o Experimental Design: To test causal relationships.
o Correlational Design: To identify relationships between variables.
7. Approaches to Research
a) Quantitative Approach
Data Type: Numerical
Methods: Surveys, experiments, statistical tests
Objective: To quantify the problem and generalize results
Example: Measuring the effect of a drug on blood pressure
b) Qualitative Approach
Data Type: Non-numerical (text, interviews, observations)
Methods: Case studies, focus groups, ethnography
Objective: To gain deep understanding of human behavior and experience
Example: Studying patient experiences during cancer treatment
8. Exploratory and Confirmatory Research
Exploratory Research:
Conducted to gain insight into an unstudied problem or identify variables; no predefined
hypothesis.
o Example: Exploring reasons for low student performance.
Confirmatory Research:
Conducted to test predefined hypotheses or theories using structured methods.
o Example: Testing whether a new teaching method improves student scores.
9. Experimental and Theoretical Research
Experimental Research:
Involves manipulation of one or more independent variables to observe the effect on
dependent variables under controlled conditions.
o Example: Testing the effect of temperature on chemical reaction rates.
Theoretical Research:
Focuses on developing abstract concepts, models, or frameworks without practical
experimentation.
o Example: Developing a mathematical model of disease spread.
Problem Formulation: Conducting Literature Review, Information Sources (Books,
monographs, reviews, blogs, etc.), Information Retrieval, Role of libraries in
Information Retrieval, Tools for identifying literature (digital resources and print media),
Indexing and abstracting services, Citation indexes, Summarizing the Review, Critical
Review, Identifying Research Gap, Conceptualizing and Hypothesizing the research
gap.
1. Problem Formulation
Meaning:
Problem formulation is the process of defining the research problem clearly and precisely.
It represents the foundation of the entire research — a poorly defined problem leads to weak
outcomes.
Key Steps:
1. Identify broad research area (e.g., AI in healthcare).
2. Review existing literature to understand what is known.
3. Narrow down to a specific issue or gap.
4. Define objectives and research questions.
5. Formulate hypotheses (if applicable).
Qualities of a Good Research Problem:
Clearly defined and specific
Researchable and feasible
Original and significant
Supported by evidence/literature
Aligned with theoretical background
Example:
Broad area – Brain tumor classification using CNNs
Problem – Low accuracy and lack of interpretability in lightweight CNN models.
2. Conducting a Literature Review
Meaning:
A literature review is a systematic, comprehensive, and critical analysis of existing research
related to a particular topic.
It helps to understand what has been done, what gaps exist, and what can be done next.
Objectives:
To understand theoretical and empirical foundations.
To identify gaps or inconsistencies in previous studies.
To avoid duplication of research.
To refine research objectives and methodology.
To build a conceptual framework for the study.
Steps in Conducting a Literature Review:
1. Define the topic or research question.
2. Search for relevant literature (journals, books, databases).
3. Screen and select studies relevant to the research problem.
4. Read critically and summarize key findings.
5. Compare and contrast different viewpoints.
6. Identify gaps and develop the research hypothesis.
3. Information Sources
A. Primary Sources
Original and first-hand materials.
Examples: Research articles, theses, patents, field data, technical reports, interviews.
B. Secondary Sources
Summaries or analyses of primary works.
Examples: Review papers, monographs, books, encyclopedias.
C. Tertiary Sources
Provide access to primary and secondary sources.
Examples: Databases, indexing and abstracting services, bibliographies.
4. Information Retrieval
Definition:
Information retrieval is the process of searching, locating, and obtaining relevant information
from various sources (digital or print) to support research.
Key Steps:
1. Define information needs (keywords, questions).
2. Select databases or repositories.
3. Use Boolean operators (AND, OR, NOT) for efficient searching.
4. Evaluate and refine results.
5. Store, organize, and cite retrieved materials.
Common Tools:
Databases: Scopus, Web of Science, IEEE Xplore, PubMed.
Search Engines: Google Scholar, Microsoft Academic.
Repositories: ResearchGate, arXiv, [Link].
5. Role of Libraries in Information Retrieval
Provide access to journals, books, and databases (both print and digital).
Offer reference services and research support.
Maintain catalogues and indexing systems for quick access.
Support Interlibrary Loan (ILL) for rare resources.
Provide training in database searching and citation management tools.
Digital Library Examples:
DELNET, NDL (National Digital Library of India), INFLIBNET, IEEE Digital Library.
6. Tools for Identifying Literature
A. Digital Resources
Online Databases: Scopus, PubMed, ScienceDirect, SpringerLink.
Search Engines: Google Scholar, Semantic Scholar.
Citation Managers: Mendeley, Zotero, EndNote.
Institutional Repositories: arXiv, ResearchGate, SSRN.
B. Print Media
Books, encyclopedias, journals, conference proceedings.
University libraries, archives, and government reports.
7. Indexing and Abstracting Services
Indexing:
A list or database of bibliographic information (title, author, keywords, source).
Helps identify relevant research quickly.
Examples: Scopus, Web of Science, MEDLINE, INSPEC.
Abstracting:
Provides summaries or abstracts of published literature.
Saves time by letting researchers decide if full text is needed.
Examples: Chemical Abstracts, Biological Abstracts, PsycINFO.
Benefits:
Enhances visibility of research.
Facilitates literature search.
Helps track citations and impact.
8. Citation Indexes
Track how many times and where a paper or author is cited.
Measure research influence and quality.
Major Citation Indexes:
Science Citation Index (SCI) – Clarivate Analytics
Social Sciences Citation Index (SSCI)
Arts and Humanities Citation Index (AHCI)
Google Scholar Citations
Use: Identify leading researchers, journals, and emerging topics.
9. Summarizing the Review
Extract key points, findings, and methodologies from each source.
Organize summaries chronologically, thematically, or methodologically.
Use tables or concept maps to present relationships among studies.
Example Table:
Author & Year Method Used Findings Gap Identified
Smith (2021) CNN-based model 90% accuracy Poor explainability
Lee (2022) XAI integration Better interpretation Limited dataset
10. Critical Review
Definition:
A critical review evaluates the strengths, weaknesses, and contributions** of existing studies.
Process:
1. Assess relevance, validity, and reliability of each study.
2. Compare different theoretical frameworks.
3. Identify contradictions, gaps, or limitations.
4. Synthesize information to form your own viewpoint.
Goal: To provide a balanced evaluation and justify the need for your study.
11. Identifying the Research Gap
Meaning:
A research gap is an unexplored or underexplored area in existing literature that needs further
study.
Ways to Identify a Gap:
Contradictory findings in previous research.
Understudied populations or contexts.
Outdated methods or technologies.
Lack of theoretical or empirical evidence.
Example:
Most studies used CNNs for brain tumor classification but few focused on explainability →
Gap: Lack of interpretable CNN models.
12. Conceptualizing and Hypothesizing the Research Gap
Conceptualization:
Process of defining and structuring the identified gap into a clear conceptual model or
framework.
Shows how variables are related and how the research will address the gap.
Hypothesizing:
Formulating testable statements (hypotheses) based on the conceptual model.
Hypothesis = “Expected relationship between variables.”
Example:
Hypothesis: Integrating explainable AI techniques into CNNs will improve diagnostic trust and
accuracy in brain tumor classification.
Research Design: Experimental / Simulation/ Theoretical /Empirical Research, Cause
effect relationship, Development of Hypothesis, Measurement Systems Analysis, Validity
and Reliability, Statistical Design of Experiments, Field Experiments, Data/Variable
Types & Classification, Data collection - Methods and Tools.
1. Research Design – Overview
Definition:
Research design is a structured plan or framework for conducting a study.
It specifies what, how, when, and where data will be collected and analyzed.
Purpose:
o To ensure valid, objective, and accurate results.
o To minimize bias and errors.
o To provide a roadmap for the entire study.
⚗️2. Types of Research Design
(a) Experimental Research
Researcher manipulates one or more independent variables and observes their effect on
dependent variables.
Objective: Establish cause–effect relationships.
Key Elements:
o Control group and experimental group.
o Random assignment.
o Independent and dependent variables.
Example:
Testing how different teaching methods affect student performance.
(b) Simulation Research
A model-based approach that imitates real-world processes through mathematical or
computer models.
Used when:
o Real experiments are costly, risky, or time-consuming.
Examples:
o Climate modeling, economic forecasting, traffic simulation.
(c) Theoretical Research
Focuses on developing new theories, models, or frameworks rather than collecting data.
Involves logical reasoning, conceptual analysis, and mathematical proofs.
Example:
Deriving a new algorithm for neural network optimization.
(d) Empirical Research
Based on observations or experiments.
Collects data from the real world to test hypotheses or theories.
Example:
Surveying 500 teachers to study job satisfaction.
🔁 3. Cause–Effect Relationship (Causality)
Definition:
A cause–effect relationship exists when a change in one variable (cause) directly leads to
a change in another variable (effect).
Conditions for Causality:
1. Temporal Precedence: Cause happens before effect.
2. Covariation: When the cause changes, the effect changes.
3. Non-spuriousness: No third variable explains the relationship.
Example:
Increasing fertilizer → increases plant growth (cause → effect).
💡 4. Development of Hypothesis
Definition:
A hypothesis is a tentative, testable statement about the relationship between two or more
variables.
Characteristics:
o Clear and specific.
o Testable and measurable.
o Based on theory or prior evidence.
Types of Hypotheses:
o Null Hypothesis (H₀): No relationship between variables.
o Alternative Hypothesis (H₁): There is a relationship.
Example:
H₀: There is no effect of study time on exam scores.
H₁: Students who study more score higher.
Steps in Developing Hypothesis:
1. Identify research problem.
2. Review existing literature.
3. Identify variables.
4. Formulate tentative relationship.
5. State hypothesis clearly.
⚙️5. Measurement Systems Analysis (MSA)
Definition:
MSA evaluates the accuracy, precision, and consistency of a measurement system.
Purpose:
To ensure that data collected are reliable and valid.
Key Terms:
o Accuracy: Closeness of measurement to true value.
o Precision: Consistency/repeatability of measurements.
o Bias: Systematic error.
o Repeatability: Same operator, same instrument, same conditions.
o Reproducibility: Different operators/instruments producing same result.
Example:
Checking consistency of temperature sensors in an experiment.
✅ 6. Validity and Reliability
(A) Validity
Definition: Degree to which an instrument measures what it is supposed to measure.
Types:
1. Content Validity: Whether the measure covers all aspects of the concept.
2. Construct Validity: Whether it measures the theoretical construct correctly.
3. Criterion Validity: Whether it correlates with other relevant measures.
Predictive Validity (future outcomes).
Concurrent Validity (current outcomes).
(B) Reliability
Definition: The consistency or stability of measurement over time.
Types:
1. Test–Retest Reliability – same test, same people, different time.
2. Inter-Rater Reliability – consistency among observers.
3. Split-Half Reliability – internal consistency of a test.
👉 Relation:
A measure must be reliable to be valid, but not all reliable measures are valid.
📈 7. Statistical Design of Experiments (DOE)
Definition:
DOE is a planned approach to conduct experiments using statistical principles to obtain
valid and reliable results.
Objectives:
o To study effects of factors (independent variables).
o To minimize experimental errors.
o To optimize performance or process.
Basic Principles (by R.A. Fisher):
1. Replication – repeating experiments to reduce random error.
2. Randomization – assigning subjects randomly to groups.
3. Local Control/Blocking – grouping similar experimental units.
Common Designs:
o Completely Randomized Design (CRD)
o Randomized Block Design (RBD)
o Factorial Design
o Latin Square Design
🧪 8. Field Experiments
Definition:
Experiments conducted in real-world settings instead of laboratories.
Advantages:
o High ecological validity.
o Realistic behavior observed.
Disadvantages:
o Less control over variables.
o Possible external interferences.
Example:
Testing a new teaching method in an actual classroom.
📊 9. Data / Variable Types & Classification
(A) Types of Data
Type Description Example
Primary Data Collected first-hand by the researcher Surveys, Experiments
Secondary Data Collected by others Census, Reports
(B) Data Based on Measurement Scale
Scale Nature Example
Nominal Categories only (no order) Gender, Religion
Ordinal Rank/order without equal intervals Satisfaction: High, Medium, Low
Interval Equal intervals, no true zero Temperature (°C)
Ratio Equal intervals with true zero Height, Weight, Income
(C) Variable Types
Type Meaning Example
Independent Variable (IV) Variable manipulated by researcher Type of fertilizer
Dependent Variable (DV) Variable measured as outcome Plant growth
Control Variable Kept constant Temperature, water level
Extraneous Variable Uncontrolled variable influencing DV Soil quality
🧰 10. Data Collection – Methods and Tools
(A) Methods
Method Description Example
Observation Watching and recording behavior Classroom observation
Survey Structured questions via questionnaire/interview Market research
Interview Direct verbal interaction In-depth interviews
Method Description Example
Experiment Controlled testing of cause–effect Drug testing
Case Study Detailed study of single case/unit Study of one company
Content Analysis Analyzing text, documents, or media Analyzing news articles
(B) Tools
Questionnaire
Interview schedule
Checklist
Rating scales (Likert, Semantic Differential)
Psychological tests
Observation sheets
Sensors/instruments (in technical research)
Data Analysis and Interpretation: Sampling, Sampling Error, Statistical
Methods/Tools - Measures of Central Tendency and Variation, Test of Hypothesis -
z test, t test, F test, ANOVA, Chi square, correlation and regression analysis, Error
Estimation.
1. Data Analysis and Interpretation – Overview
Definition:
Data Analysis is the systematic process of inspecting, cleaning, transforming, and
modeling data to discover useful information and reach conclusions.
Interpretation involves drawing logical conclusions from analyzed data.
Purpose:
o To test hypotheses.
o To identify patterns or relationships.
o To derive insights and make decisions.
🧮 2. Sampling
(a) Definition:
Sampling is the process of selecting a subset of individuals (sample) from a population to draw
conclusions about the whole.
(b) Population vs Sample
Term Meaning
Population Entire group under study
Sample A small part representing the population
(c) Types of Sampling
I. Probability Sampling (Each unit has equal chance)
Method Description
Simple Random Sampling Every member has an equal chance (e.g., lottery method).
Systematic Sampling Every kth element from a list.
Stratified Sampling Population divided into strata (groups) → random sample from each.
Cluster Sampling Population divided into clusters → randomly select some clusters.
II. Non-Probability Sampling
Method Description
Convenience Sampling Based on availability/ease.
Judgmental/Purposive Sampling Researcher’s expert judgment.
Quota Sampling Sample according to pre-set quotas.
Existing subjects recruit future subjects (used in social
Snowball Sampling
research).
(d) Sampling Error
Definition: Difference between sample statistic and true population parameter due to
using a sample instead of the population.
Types:
1. Random Sampling Error: Due to chance variation.
2. Systematic Error (Bias): Due to improper sampling methods.
Formula for Standard Error (SE):
SE=snSE = \frac{s}{\sqrt{n}}SE=ns
where s = sample standard deviation, n = sample size.
📈 3. Statistical Methods and Tools
(a) Measures of Central Tendency
Measure Formula Example
Mean (Arithmetic Average) Xˉ=∑XN\bar{X} = \frac{\sum X}{N}Xˉ=N∑X
Median (Middle Value) Middle of ordered data
Mode (Most Frequent Value) Value occurring most often
Relation:
For a normal distribution → Mean = Median = Mode
(b) Measures of Variation (Dispersion)
Measure Meaning Formula
R=Xmax−XminR = X_{max} - X_{min}R=Xmax
Range Max – Min
−Xmin
Average of squared σ2=∑(X−Xˉ)2N\sigma^2 = \frac{\sum (X - \
Variance
deviations bar{X})^2}{N}σ2=N∑(X−Xˉ)2
Standard Deviation Square root of σ=∑(X−Xˉ)2N\sigma = \sqrt{\frac{\sum (X - \
(SD) variance bar{X})^2}{N}}σ=N∑(X−Xˉ)2
Coefficient of Relative measure of CV=SDMean×100CV = \frac{SD}{Mean} \times
Variation (CV) spread 100CV=MeanSD×100
📊 4. Test of Hypothesis
(a) Definition:
A statistical test used to decide whether to accept or reject a null hypothesis (H₀) based on
sample data.
(b) Steps in Hypothesis Testing:
1. Formulate H₀ (null) and H₁ (alternative).
2. Select significance level (α) (commonly 0.05 or 0.01).
3. Choose appropriate statistical test (z, t, F, etc.).
4. Compute test statistic.
5. Compare with critical value or p-value.
6. Accept/Reject H₀.
(c) Types of Errors
Type Meaning
Type I Error (α) Rejecting true H₀ (False positive)
Type II Error (β) Accepting false H₀ (False negative)
🔢 5. Major Statistical Tests
(a) z-Test
Used for: Large samples (n > 30)
Purpose: Test difference between means or proportions when σ (population SD) is
known.
Formula:
z=Xˉ−μσ/nz = \frac{\bar{X} - \mu}{\sigma / \sqrt{n}}z=σ/nXˉ−μ
Example: Checking if sample mean differs from population mean.
(b) t-Test
Used for: Small samples (n < 30)
Population SD unknown.
Types:
1. One-sample t-test: Compare sample mean with population mean.
2. Independent t-test: Compare means of two independent groups.
3. Paired t-test: Compare means of the same group before & after treatment.
t=Xˉ1−Xˉ2Sp1n1+1n2t = \frac{\bar{X}_1 - \bar{X}_2}{S_p \sqrt{\frac{1}{n_1} + \frac{1}
{n_2}}}t=Spn11+n21Xˉ1−Xˉ2
where SpS_pSp = pooled SD.
(c) F-Test
Purpose: Compare two variances.
F=S12S22F = \frac{S_1^2}{S_2^2}F=S22S12
Basis of ANOVA (Analysis of Variance).
(d) ANOVA (Analysis of Variance)
Used to: Compare three or more group means simultaneously.
Logic: Total variance is divided into variance between groups and within groups.
F=Between-group varianceWithin-group varianceF = \frac{\text{Between-group variance}}{\
text{Within-group variance}}F=Within-group varianceBetween-group variance
Types:
o One-way ANOVA (1 factor)
o Two-way ANOVA (2 factors)
(e) Chi-Square (χ²) Test
Used for: Testing relationships between categorical variables.
Formula:
χ2=∑(O−E)2E\chi^2 = \sum \frac{(O - E)^2}{E}χ2=∑E(O−E)2
where O = observed frequency, E = expected frequency.
Applications:
Test of independence.
Goodness of fit test.
(f) Correlation Analysis
Measures: Degree of association between two variables.
Types:
o Positive correlation (+)
o Negative correlation (–)
o Zero correlation (no relationship)
Pearson’s correlation coefficient (r):
r=N∑XY−∑X∑Y[N∑X2−(∑X)2][N∑Y2−(∑Y)2]r = \frac{N\sum XY - \sum X \sum Y}{\
sqrt{[N\sum X^2 - (\sum X)^2][N\sum Y^2 - (\sum Y)^2]}}r=[N∑X2−(∑X)2][N∑Y2−(∑Y)2]
N∑XY−∑X∑Y
Range: –1 ≤ r ≤ +1
Interpretation:
r = +1 → Perfect positive
r = –1 → Perfect negative
r = 0 → No correlation
(g) Regression Analysis
Purpose: To predict the value of one variable based on another.
Simple Linear Regression Equation:
Y=a+bXY = a + bXY=a+bX
where
b = regression coefficient = Cov(X,Y)Var(X)\frac{\text{Cov}(X,Y)}{\text{Var}
(X)}Var(X)Cov(X,Y),
a = intercept.
Interpretation:
Shows direction and strength of relationship between independent (X) and dependent (Y)
variables.
⚙️6. Error Estimation
Definition:
Process of determining the accuracy and precision of results obtained from sample data.
Types of Errors:
1. Sampling Error: Due to sample not representing the population.
2. Non-Sampling Error: Due to data collection, recording, or analysis mistakes.
3. Measurement Error: Due to faulty instruments or respondent bias.
Error Reduction Methods:
Increase sample size.
Use proper sampling techniques.
Calibrate instruments.
Train data collectors.
Double-check data entry.
📘 7. Data Interpretation
Definition:
Process of making sense of analyzed data, identifying patterns, relationships, and drawing
conclusions.
Steps:
1. Summarize data findings.
2. Compare with hypotheses.
3. Draw inferences or implications.
4. Validate with literature or theory.
5. Report results clearly (tables, graphs, charts).
🧠 8. Summary for Quick Revision
Concept Key Points
Sampling Selects representative part of population
Sampling Error Difference between sample and population
Central Tendency Mean, Median, Mode
Dispersion Range, Variance, SD, CV
Hypothesis Tests z, t, F, ANOVA, χ²
Correlation Strength of relationship (r)
Regression Predicts Y from X
Validity Accuracy
Reliability Consistency
Error Estimation Evaluates accuracy and precision of results
Writing Research Articles and Thesis: Data Presentation - Types of tables and
illustrations, Guidelines for writing the abstract, introduction, methodology, results
and discussion, conclusion sections of a manuscript. References - Styles and
methods, Citation and listing system of documents. Plagiarism. Ethical considerations in
Research
1. Writing Research Articles and Thesis
Definition
Research writing involves communicating the outcomes of scientific inquiry in a clear,
structured, and logical form — usually as a research paper, dissertation, or thesis.
Purpose:
To share findings, contribute to existing knowledge, and allow replication by others.
📊 2. Data Presentation
Meaning:
Data presentation refers to organizing analyzed data in a meaningful, concise, and visual way for
better understanding.
(A) Types of Data Presentation
Type Form Examples
Data explained in words or “Out of 100 respondents, 60%
Textual
sentences preferred online classes.”
Tabular Data arranged in rows & columns Frequency tables, comparison tables
Graphical / Bar graphs, pie charts, line graphs,
Data shown pictorially
Diagrammatic histograms
Visuals, maps, photographs, Conceptual frameworks, images of
Illustrative
flowcharts, models apparatus
(B) Types of Tables and Illustrations
Type Purpose
Simple Table Displays one variable (e.g., gender distribution)
Complex Table Displays more than one variable (e.g., gender × qualification)
Frequency Table Shows frequency counts
Cross-tabulation For comparing multiple variables
Illustrations Diagrams, charts, graphs, maps, flowcharts for clarity
Tips:
Number tables and figures separately (Table 1, Figure 1).
Provide clear titles and units.
Mention source below if data is secondary.
Refer to them properly in text (“As shown in Table 2…”).
✍️3. Structure of a Research Paper / Thesis
A standard research article or thesis follows IMRaD format:
Section Purpose Key Points
Title Concise & informative Should reflect key variables & design
150–250 words; brief intro, method, results,
Abstract Summary of study
conclusion
Keywords 4–6 important words Help in indexing/search
Introduction Background & rationale Problem, objectives, significance, literature gap
How research was Design, sample, instruments, procedure, data
Methodology
conducted analysis
Results Findings of the study Tables, figures, statistical results (no interpretation)
Discussion Interpretation of results Explain implications, compare with past studies
Conclusion Final summary & outcomes Key findings, limitations, future scope
References List of cited works Follow standard format (APA, IEEE, etc.)
Appendices Supplementary material Questionnaire, raw data, codes, etc.
📄 4. Section-wise Writing Guidelines
(A) Abstract
Purpose: Summarize the entire paper.
Includes:
o Purpose of study
o Method used
o Major findings
o Conclusions/implications
Tips:
o Write last, though placed first.
o Avoid citations and abbreviations.
o Limit 150–250 words.
(B) Introduction
Purpose: To introduce topic and justify the need for study.
Includes:
1. Background and context.
2. Statement of the problem.
3. Research objectives and questions.
4. Hypotheses (if applicable).
5. Scope and significance.
(C) Methodology
Purpose: To describe how the research was conducted.
Includes:
o Research design (qualitative/quantitative/mixed).
o Sampling method and size.
o Tools/instruments used (e.g., questionnaire, sensor).
o Data collection procedure.
o Statistical methods used for analysis.
o Ethical considerations.
(D) Results
Purpose: Present data findings.
Includes:
o Tables, graphs, charts.
o Statistical test results (mean, SD, t, F, χ², correlation, etc.).
o Avoid interpretation — only state findings.
(E) Discussion
Purpose: Explain and interpret results.
Includes:
o Comparison with previous studies.
o Theoretical implications.
o Practical applications.
o Reasons for unexpected results.
(F) Conclusion
Purpose: Provide final insights.
Includes:
o Major findings.
o Contribution to knowledge.
o Limitations.
o Suggestions for future research.
📚 5. References – Styles and Methods
(A) Meaning:
References acknowledge the sources used in your research work.
(B) Importance:
Avoids plagiarism.
Adds credibility.
Allows readers to verify sources.
(C) Major Reference Styles
In-text
Style Used In Reference List Example
Citation
APA (American Smith, J. (2022). Title of
Social Sciences (Smith, 2022)
Psychological Association) book. Publisher.
MLA (Modern Language Smith, John. Title of Book.
Humanities (Smith 22)
Association) Publisher, 2022.
Superscript Smith, John. Title of Book.
Chicago History, Arts
numbers Chicago Press, 2022.
Smith, J. (2022). Title of
Harvard General (Smith, 2022)
book. London: Publisher.
Engineering & [1] J. Smith, Title of Book,
IEEE [1], [2]
Technology Publisher, 2022.
(D) Citation and Listing System
1. Citation:
Mention within the text to indicate source.
o Example (APA): “...as explained by Johnson (2021).”
o Example (IEEE): “...as reported in [3].”
2. Listing (Bibliography/References):
Full details of sources at end of document.
o Arrange alphabetically (APA/Harvard) or numerically (IEEE).
🚫 6. Plagiarism
Definition:
Plagiarism is the use of another person’s ideas, words, or work without proper acknowledgment.
Types:
1. Direct plagiarism: Copying text word-for-word.
2. Mosaic plagiarism: Mixing copied material with own words.
3. Self-plagiarism: Reusing one’s previous work.
4. Accidental plagiarism: Failing to cite properly.
Prevention:
Use quotation marks for direct quotes.
Cite every source accurately.
Use plagiarism checkers (Turnitin, iThenticate, Grammarly).
Paraphrase properly with source credit.
⚖️7. Ethical Considerations in Research
Aspect Guidelines
Participants must be informed about the study’s purpose and give
Informed Consent
consent.
Confidentiality Protect participants’ identities and data.
Honesty & Integrity Report data truthfully, avoid fabrication.
Avoid Harm Protect participants from physical or psychological harm.
Respect Intellectual
Acknowledge sources and contributions.
Property
Transparency Disclose funding sources, conflicts of interest.
Ethical Review:
Before starting, research should be approved by an Ethical Review Committee or Institutional
Review Board (IRB).
🧠 8. Summary for Quick Revision
Concept Key Points
Data Presentation Tables, graphs, diagrams – clear, labeled, referenced
Abstract 150–250 words; summary of aim, method, results, conclusion
Introduction Background, problem, objectives
Methodology Design, sample, tools, data collection, analysis
Concept Key Points
Results & Discussion Findings → Interpretations & implications
Conclusion Key insights, limitations, future work
References Follow APA/IEEE/Harvard; cite properly
Plagiarism Unauthorized use of others’ work
Ethics Consent, confidentiality, honesty, respect
UNIT I Probability and Statistics
Random Variables, Probability Distributions, Correlation, Regression, Testing of hypothesis;
1. Random Variables
➤ Definition:
A random variable (RV) is a variable that takes numerical values determined by the outcome of
a random experiment.
Types:
1. Discrete Random Variable:
Takes countable values.
Example: No. of heads in 3 coin tosses → {0,1,2,3}
2. Continuous Random Variable:
Takes uncountable values within a range.
Example: Height, weight, temperature.
➤ Probability Distribution Function (PDF/PMF)
For discrete RV:
P(X=xi)=piP(X = x_i) = p_iP(X=xi)=pi
where ∑pi=1\sum p_i = 1∑pi=1
For continuous RV:
f(x)≥0f(x) \ge 0f(x)≥0 and ∫−∞∞f(x) dx=1\int_{-\infty}^{\infty} f(x)\,dx = 1∫−∞∞
f(x)dx=1
Cumulative Distribution Function (CDF):
F(x)=P(X≤x)F(x) = P(X \le x)F(x)=P(X≤x)
➤ Expected Value and Variance
Mean (Expectation):
o Discrete: E(X)=∑xipiE(X) = \sum x_i p_iE(X)=∑xipi
o Continuous: E(X)=∫xf(x) dxE(X) = \int x f(x)\,dxE(X)=∫xf(x)dx
Variance:
Var(X)=E(X2)−[E(X)]2Var(X) = E(X^2) - [E(X)]^2Var(X)=E(X2)−[E(X)]2
🔹 2. Probability Distributions
(A) Discrete Distributions
Distribution Parameters Mean Variance PMF / Notes
Bernoulli p p p(1–p) One trial: success/failure
P(X=k)=(nk)pk(1−p)n−kP(X=k)=\binom{n}
Binomial n, p np np(1–p)
{k}p^k(1-p)^{n-k}P(X=k)=(kn)pk(1−p)n−k
P(X=k)=e−λλkk!P(X=k)=\frac{e^{-\lambda}\
Poisson λ λ λ
lambda^k}{k!}P(X=k)=k!e−λλk (rare events)
Geometric p 1/p (1–p)/p² Trials till first success
(B) Continuous Distributions
Distribution Parameters Mean Variance PDF / Notes
Uniform a, b (a+b)/2 (b–a)²/12 Constant probability
f(x)=12πσ2e−(x−μ)22σ2f(x)=\frac{1}{\
Normal
μ, σ μ σ² sqrt{2πσ^2}} e^{-\frac{(x-μ)^2}{2σ^2}}f(x)=2πσ2
(Gaussian)
1e−2σ2(x−μ)2
Exponential λ 1/λ 1/λ² f(x)=λe−λxf(x)=λe^{-λx}f(x)=λe−λx, memoryless
Chi-Square k k 2k Used in hypothesis tests
t-
df 0 >1 Small samples (<30)
Distribution
Distribution Parameters Mean Variance PDF / Notes
F-
df₁, df₂ — — Ratio of variances
Distribution
🔹 3. Correlation
➤ Definition:
Measures the strength and direction of the linear relationship between two variables (X and Y).
➤ Formula (Karl Pearson’s Correlation Coefficient):
r=n∑xy−(∑x)(∑y)[n∑x2−(∑x)2][n∑y2−(∑y)2]r = \frac{n\sum xy - (\sum x)(\sum y)}{\sqrt{[n\
sum x^2 - (\sum x)^2][n\sum y^2 - (\sum y)^2]}}r=[n∑x2−(∑x)2][n∑y2−(∑y)2]n∑xy−(∑x)(∑y)
Range: –1 ≤ r ≤ +1
o r = +1 → Perfect positive correlation
o r = –1 → Perfect negative correlation
o r = 0 → No correlation
➤ Spearman’s Rank Correlation:
Used for ranked or non-parametric data.
rs=1−6∑di2n(n2−1)r_s = 1 - \frac{6\sum d_i^2}{n(n^2 - 1)}rs=1−n(n2−1)6∑di2
🔹 4. Regression Analysis
➤ Definition:
Regression estimates the relationship between a dependent variable (Y) and one or more
independent variables (X).
➤ Simple Linear Regression:
Y=a+bXY = a + bXY=a+bX
Where:
b=n∑xy−(∑x)(∑y)n∑x2−(∑x)2b = \frac{n\sum xy - (\sum x)(\sum y)}{n\sum x^2 - (\sum
x)^2}b=n∑x2−(∑x)2n∑xy−(∑x)(∑y)
a=Yˉ−bXˉa = \bar{Y} - b\bar{X}a=Yˉ−bXˉ
Purpose: Predict value of Y for given X.
Coefficient of Determination (R²):
R2=r2R^2 = r^2R2=r2
Shows percentage of variance explained by regression.
🔹 5. Testing of Hypothesis
➤ Meaning:
A statistical method to test an assumption (hypothesis) about a population using sample data.
➤ Steps:
1. State Null (H₀) and Alternative (H₁) hypotheses
2. Choose Significance Level (α) (e.g., 0.05 or 5%)
3. Select appropriate Test Statistic
4. Compute the statistic from data
5. Compare with critical value or p-value
6. Accept or reject H₀
➤ Common Tests:
Test Use Formula / Condition
Large samples (n > 30), z=Xˉ−μσ/nz = \frac{\bar{X} - μ}{σ/\
Z-Test
known σ sqrt{n}}z=σ/nXˉ−μ
Small samples (n < 30), t=Xˉ−μs/nt = \frac{\bar{X} - μ}{s/\
t-Test
unknown σ sqrt{n}}t=s/nXˉ−μ
Compare variances of two
F-Test F=s12s22F = \frac{s_1^2}{s_2^2}F=s22s12
samples
ANOVA (Analysis of Compare means of ≥ 3
Based on F-distribution
Variance) groups
Test of independence or χ2=∑(O−E)2Eχ^2 = \sum \frac{(O - E)^2}
Chi-Square Test
goodness of fit {E}χ2=∑E(O−E)2
➤ Errors in Hypothesis Testing:
Type Description
Type I Error (α) Rejecting a true null hypothesis
Type II Error (β) Accepting a false null hypothesis
Power of Test 1 – β (probability of correctly rejecting false H₀)
🔹 6. Error Estimation
➤ Standard Error (SE):
Measures variability between sample mean and population mean.
SE=snSE = \frac{s}{\sqrt{n}}SE=ns
➤ Confidence Interval (CI):
Range within which population parameter likely lies.
For mean:
Xˉ±Zα/2×SE\bar{X} \pm Z_{\alpha/2} \times SEXˉ±Zα/2×SE
Example: 95% CI → Z = 1.96
🔹 7. Key Formulas Summary
Concept Formula
Mean Xˉ=∑Xin\bar{X} = \frac{\sum X_i}{n}Xˉ=n∑Xi
Variance σ2=∑(Xi−Xˉ)2nσ^2 = \frac{\sum (X_i - \bar{X})^2}{n}σ2=n∑(Xi−Xˉ)2
Standard
σ=σ2σ = \sqrt{σ^2}σ=σ2
Deviation
Cov(X,Y)=∑(Xi−Xˉ)(Yi−Yˉ)n−1Cov(X,Y) = \frac{\sum (X_i - \bar{X})(Y_i - \
Covariance
bar{Y})}{n-1}Cov(X,Y)=n−1∑(Xi−Xˉ)(Yi−Yˉ)
Correlation r=Cov(X,Y)σXσYr = \frac{Cov(X,Y)}{σ_X σ_Y}r=σXσYCov(X,Y)
Regression (Y
Y=a+bXY = a + bXY=a+bX
on X)
Concept Formula
Z-statistic z=Xˉ−μσ/nz = \frac{\bar{X} - μ}{σ/\sqrt{n}}z=σ/nXˉ−μ
t-statistic t=Xˉ−μs/nt = \frac{\bar{X} - μ}{s/\sqrt{n}}t=s/nXˉ−μ
F-statistic F=s12s22F = \frac{s_1^2}{s_2^2}F=s22s12
Chi-Square χ2=∑(O−E)2Eχ^2 = \sum \frac{(O - E)^2}{E}χ2=∑E(O−E)2
🔹 8. Example Questions (PhD Entrance Pattern)
1. If X ~ Binomial (n = 10, p = 0.4), find mean and variance.
→ Mean = 10×0.4 = 4; Variance = 10×0.4×0.6 = 2.4
2. What is the probability of getting 0 defects if average defects λ = 2 (Poisson)?
→ P(X=0)=e−2(20)/0!=e−2=0.1353P(X=0) = e^{-2} (2^0)/0! = e^{-2} =
0.1353P(X=0)=e−2(20)/0!=e−2=0.1353
3. Two samples have variances 10 and 15; test equality using F-test.
→ F = 15/10 = 1.5 (compare with critical F at given df)
4. If r = 0.9, what % of variance in Y is explained by X?
→ R2=0.92=0.81=81%R^2 = 0.9^2 = 0.81 = 81\%R2=0.92=0.81=81%
🔹 9. Quick Concept Recap Table
Concept Key Idea Common Tool
Random Variable Outcome → numeric value PMF/PDF
Distribution Probability model Binomial, Normal
Correlation Strength of relation r
Regression Prediction model Y = a + bX
Hypothesis Test assumption Z, t, F, χ²
Error Estimation Sample accuracy SE, CI
Theory of Computation - Finite State Machine, Pushdown Automata, Context Free
Grammar, Turing Machine
1. Introduction to Theory of Computation
Definition:
Theory of Computation studies how problems can be solved on a model of computation (like
automata or Turing machines) and the limits of what can be computed.
Main Areas:
1. Automata Theory – Machine models for computation (FSM, PDA, TM)
2. Formal Languages – Mathematical description of languages
3. Computability – What problems can/cannot be solved
4. Complexity Theory – How efficiently problems can be solved
🔹 2. Finite State Machine (Finite Automata)
➤ Definition:
A Finite State Machine (FSM) is a mathematical model of computation consisting of a finite
number of states. It recognizes Regular Languages.
➤ Components of FSM:
An FSM is represented as a 5-tuple:
M=(Q,Σ,δ,q0,F)M = (Q, \Sigma, \delta, q_0, F)M=(Q,Σ,δ,q0,F)
where:
Q → Finite set of states
Σ → Input alphabet
δ → Transition function (δ: Q × Σ → Q)
q₀ → Start state
F → Set of final (accepting) states
➤ Types:
1. Deterministic Finite Automata (DFA):
o One transition per symbol per state
o No ε-moves
o Easy to implement
2. Non-Deterministic Finite Automata (NFA):
o Can have multiple transitions for same input
o May include ε-transitions
o Equivalent to DFA in power
➤ Example (DFA):
Language: strings over {0,1} ending with 1
States: q₀ (start), q₁ (accept)
Transitions:
q₀ → 0 → q₀
q₀ → 1 → q₁
q₁ → 0 → q₀
q₁ → 1 → q₁
➤ Regular Expressions and FSM:
Regular expressions (RE) describe Regular Languages, which are exactly those recognized by
FSMs.
Examples:
(a+b)* → all strings over {a,b}
a*b → any number of a’s followed by one b
➤ Limitations of FSM:
Cannot count or remember unbounded input (e.g., {aⁿbⁿ})
Hence cannot recognize non-regular languages.
🔹 3. Pushdown Automata (PDA)
➤ Definition:
A Pushdown Automaton is a finite automaton equipped with an auxiliary stack memory. It
recognizes Context-Free Languages (CFLs).
➤ Components:
PDA is a 7-tuple:
M=(Q,Σ,Γ,δ,q0,Z0,F)M = (Q, \Sigma, \Gamma, \delta, q_0, Z_0, F)M=(Q,Σ,Γ,δ,q0,Z0,F)
where:
Q → Finite set of states
Σ → Input alphabet
Γ → Stack alphabet
δ → Transition function (δ: Q × (Σ ∪ {ε}) × Γ → P(Q × Γ*))
q₀ → Start state
Z₀ → Start symbol on stack
F → Set of final states
➤ Working Principle:
Reads input symbol
Checks stack top
Pushes or pops symbols based on transition rules
➤ Example:
Language L = {aⁿbⁿ | n ≥ 0}
Operation:
1. Push each ‘a’ onto stack.
2. For each ‘b’, pop one ‘a’.
3. Accept if stack is empty and input ends.
➤ Acceptance Criteria:
1. Final state acceptance, or
2. Empty stack acceptance
➤ Power:
PDA > FSM
PDA can recognize nested and recursive patterns (like parentheses).
🔹 4. Context-Free Grammar (CFG)
➤ Definition:
A Context-Free Grammar (CFG) is a set of production rules that describe Context-Free
Languages (CFLs).
Used in compiler design and syntax analysis.
➤ Components:
A CFG is a 4-tuple:
G=(V,Σ,R,S)G = (V, \Sigma, R, S)G=(V,Σ,R,S)
where:
V → Non-terminal symbols
Σ → Terminal symbols (alphabet)
R → Production rules (A → α)
S → Start symbol
➤ Example:
L = {aⁿbⁿ | n ≥ 1}
Grammar:
S → aSb | ab
➤ Derivations:
Leftmost derivation: Expand leftmost nonterminal first.
Rightmost derivation: Expand rightmost nonterminal first.
➤ Parse Tree:
Shows hierarchical structure of derivation.
➤ Ambiguity:
If a grammar produces more than one parse tree for the same string → ambiguous grammar.
➤ Chomsky Hierarchy of Grammars:
Type Grammar Class Recognized By Example Language
Type-0 Unrestricted Turing Machine All computable languages
Type-1 Context-Sensitive Linear Bounded Automaton {aⁿbⁿcⁿ}
Type-2 Context-Free Pushdown Automaton {aⁿbⁿ}
Type-3 Regular Finite Automata (a+b)*
🔹 5. Turing Machine (TM)
➤ Definition:
A Turing Machine (TM) is the most powerful computational model capable of simulating any
algorithm.
It recognizes Recursively Enumerable Languages (Type-0).
➤ Components:
A TM is a 7-tuple:
M=(Q,Σ,Γ,δ,q0,qaccept,qreject)M = (Q, \Sigma, \Gamma, \delta, q_0, q_{accept},
q_{reject})M=(Q,Σ,Γ,δ,q0,qaccept,qreject)
where:
Q → Finite set of states
Σ → Input alphabet
Γ → Tape alphabet (includes blank symbol ☐)
δ → Transition function: Q × Γ → Q × Γ × {L, R}
q₀ → Start state
q_accept → Accept state
q_reject → Reject state
➤ Working:
The tape is infinite and acts as memory.
Head reads and writes symbols, moves Left (L) or Right (R).
Computation halts when accept/reject state is reached.
➤ Example:
Language L = {aⁿbⁿ | n ≥ 1}
TM steps:
1. Replace leftmost ‘a’ by X, find matching ‘b’, replace with Y.
2. Repeat until no unmatched symbols remain.
3. Accept if all symbols are X/Y.
➤ Variants of TM:
Multi-tape TM – multiple tapes (more efficient, same power)
Non-deterministic TM – parallel branches (same power as deterministic)
Universal TM – can simulate any TM (basis of real computers)
➤ Halting Problem:
Some problems cannot be decided by any TM (non-computable).
Example: Determining if a TM halts on all inputs → undecidable problem.
🔹 6. Relationship Between Machines
Model Recognizes Power Example
FSM Regular Languages Lowest (a+b)*
PDA Context-Free Languages Medium {aⁿbⁿ}
TM Recursively Enumerable Highest All computable problems
🔹 7. Comparison Table
Feature FSM PDA TM
Memory Finite Stack (LIFO) Infinite tape
Language Type Regular Context-Free Recursively Enumerable
Acceptance Final state Final/Empty stack Accept/Reject states
Feature FSM PDA TM
Power Lowest Medium Highest
Example Language (a+b)* {aⁿbⁿ} {aⁿbⁿcⁿ}
🔹 8. Example Exam-Type Questions
1. Q: Which of the following languages is not regular?
A. (a+b)*
B. aⁿbⁿ
C. ab
D. (ab)*
→ Ans: B (requires counting → non-regular)
2. Q: PDA is used to recognize which type of languages?
→ Context-Free Languages
3. Q: Which is the most powerful computational model?
→ Turing Machine
4. Q: Grammar S → aSb | ab represents which language?
→ {aⁿbⁿ | n ≥ 1}
5. Q: Halting problem is ________.
→ Undecidable
🔹 9. Quick Revision Points
FSM = Regular Languages = RE equivalent
PDA = Context-Free Languages = Stack memory
TM = Universal computation = Any computable function
Hierarchy: Regular ⊂ CFL ⊂ CSL ⊂ Recursive ⊂ Recursively Enumerable
CFG → PDA → TM form an ascending chain of computational power.
UNIT II Data Structures and Algorithms
Arrays, Lists, Stacks, Queues, Trees, Graphs, Searching and Sorting Algorithms;
1. Introduction to Data Structures
Definition:
A data structure is a systematic way of organizing, managing, and storing data so that it
can be accessed and modified efficiently.
Types of Data Structures:
1. Primitive: int, float, char, double, etc.
2. Non-Primitive: Arrays, Lists, Stacks, Queues, Trees, Graphs, etc.
3. Linear: Data arranged in a sequence (e.g., Array, Linked List).
4. Non-linear: Data arranged hierarchically (e.g., Tree, Graph).
🔹 2. Arrays
Definition:
A collection of elements of the same data type stored in contiguous memory locations.
Advantages:
o Easy access (via index)
o Efficient for sequential access
Disadvantages:
o Fixed size
o Insertion/deletion costly
Operations:
o Traversing, Searching, Insertion, Deletion, Sorting
Complexities:
o Access: O(1)
o Search: O(n)
o Insert/Delete: O(n)
🔹 3. Lists (Linked Lists)
Definition:
A dynamic data structure made of nodes, each containing data + pointer to the next
node.
Types:
1. Singly Linked List – one direction link
2. Doubly Linked List – forward & backward link
3. Circular Linked List – last node links back to first node
Advantages:
o Dynamic size
o Easy insertion/deletion
Disadvantages:
o Sequential access only (no random access)
o Extra memory for pointers
Time Complexities:
o Access/Search: O(n)
o Insert/Delete: O(1) (if position known)
🔹 4. Stacks
Definition:
A LIFO (Last In, First Out) data structure.
Operations:
o Push(x): Add item to top
o Pop(): Remove item from top
o Peek(): View top item
Applications:
o Function calls, recursion
o Expression evaluation (prefix/infix/postfix)
o Undo/Redo operations
Implementation:
o Array or Linked List
Complexity:
o Push/Pop/Peek → O(1)
🔹 5. Queues
Definition:
A FIFO (First In, First Out) data structure.
Operations:
o Enqueue(x): Insert at rear
o Dequeue(): Remove from front
Types:
1. Simple Queue
2. Circular Queue
3. Priority Queue
4. Deque (Double Ended Queue)
Applications:
o CPU Scheduling, Buffer management, Printing queues
Complexity:
o Enqueue/Dequeue → O(1)
🔹 6. Trees
Definition:
A hierarchical non-linear data structure consisting of nodes connected by edges.
Terminology:
o Root: Topmost node
o Leaf: Node with no children
o Height: Longest path from root to leaf
o Degree: Number of children a node has
Types of Trees:
1. Binary Tree: Each node ≤ 2 children
2. Full Binary Tree: Every node has 0 or 2 children
3. Complete Binary Tree: All levels filled except possibly last
4. Binary Search Tree (BST): Left < Root < Right
5. AVL Tree: Self-balancing BST
6. Heap Tree: Complete binary tree satisfying heap property
Max-Heap: Parent ≥ children
Min-Heap: Parent ≤ children
7. B-Tree / B+ Tree: Used in databases and file systems
Tree Traversals:
o Inorder (LNR) – Left, Node, Right
o Preorder (NLR) – Node, Left, Right
o Postorder (LRN) – Left, Right, Node
o Level Order – BFS traversal
Complexities (BST):
o Search/Insert/Delete:
O(log n) average
O(n) worst (unbalanced tree)
🔹 7. Graphs
Definition:
A collection of nodes (vertices) and edges (links) connecting pairs of nodes.
Types:
o Directed / Undirected
o Weighted / Unweighted
o Cyclic / Acyclic
o Connected / Disconnected
Representations:
1. Adjacency Matrix → O(V²) space
2. Adjacency List → O(V + E) space
Graph Traversal Algorithms:
o BFS (Breadth-First Search): Uses Queue
o DFS (Depth-First Search): Uses Stack / Recursion
Applications:
o Network routing, Social networks, Path finding (Dijkstra, A*), Scheduling
🔹 8. Searching Algorithms
Algorithm Description Complexity
Linear Search Sequentially checks each element O(n)
Binary Search Divides sorted array into halves O(log n)
Hashing Direct access via hash function O(1) average
Binary Search Requirements: Array must be sorted.
🔹 9. Sorting Algorithms
Algorithm Best Average Worst Stable? Remarks
Bubble Sort O(n) O(n²) O(n²) Yes Simple but slow
Selection Sort O(n²) O(n²) O(n²) No Minimum element selected each pass
Insertion Sort O(n) O(n²) O(n²) Yes Efficient for small/near-sorted data
Merge Sort O(n log n) O(n log n) O(n log n) Yes Divide & Conquer
Quick Sort O(n log n) O(n log n) O(n²) No Pivot-based; fast on average
Heap Sort O(n log n) O(n log n) O(n log n) No Based on Heap tree
Radix Sort O(nk) O(nk) O(nk) Yes Non-comparative (digit-wise)
Algorithm Best Average Worst Stable? Remarks
🔹 10. Algorithm Analysis
Complexity Types:
o Time Complexity: Execution time based on input size.
o Space Complexity: Memory usage.
Common Growth Orders:
o O(1) < O(log n) < O(n) < O(n log n) < O(n²) < O(2ⁿ) < O(n!)
Asymptotic Notations:
o Big O (O): Upper bound
o Omega (Ω): Lower bound
o Theta (Θ): Tight bound
🔹 11. Sampling Concepts in Algorithms
In research-oriented algorithmic studies:
Input Sampling: Used to test performance under random data.
Error Estimation: Difference between theoretical and observed performance.
🔹 12. Applications of Data Structures
Structure Applications
Array Lookup tables, matrices
Linked List Dynamic memory allocation
Stack Recursion, expression parsing
Queue Process scheduling
Tree Databases, XML parsing
Structure Applications
Graph Network routing, dependency graphs
🧾 Quick Revision Summary
Array = Static linear storage
Linked List = Dynamic linear storage
Stack = LIFO
Queue = FIFO
Tree = Hierarchical (BST, AVL, Heap)
Graph = Network (BFS, DFS)
Sorting = Bubble, Selection, Insertion, Merge, Quick, Heap
Searching = Linear, Binary, Hashing
Complexities: Know O(1), O(n), O(log n), O(n log n), O(n²)
Programming - C, Object Oriented Programming
Part A: Programming in C
🔹 1. Overview of C Language
Developed by: Dennis Ritchie (AT&T Bell Labs, 1972)
Type: Procedural, structured, compiled language
Features: Portability, low-level access, modularity, speed, pointer manipulation
🔹 2. Structure of a C Program
#include <stdio.h>
int main() {
printf("Hello, World!");
return 0;
}
Explanation:
#include <stdio.h> → Preprocessor directive
main() → Entry point of program
printf() → Standard output function
return 0; → Successful termination
🔹 3. Data Types and Variables
Basic Types:
int, float, char, double
Derived Types:
arrays, pointers, structures, unions, functions
Example:
int a = 10;
float b = 3.14;
char c = 'A';
Storage Classes:
auto, register, static, extern
🔹 4. Operators
Category Operators Example
Arithmetic +, -, *, /, % a+b
Relational >, <, >=, <=, ==, != a>b
Logical &&, ||, ! a && b
Bitwise &, |, ^, ~, <<, >> a << 2
Assignment =, +=, -=, *= a+=2
Increment/Decrement ++, -- i++
🔹 5. Control Structures
Conditional:
if (x > 0) printf("Positive");
else if (x < 0) printf("Negative");
else printf("Zero");
Switch Case:
switch(choice) {
case 1: printf("Yes"); break;
case 2: printf("No"); break;
default: printf("Invalid");
}
Loops:
for, while, do-while
for(int i=0; i<5; i++)
printf("%d ", i);
🔹 6. Functions
Syntax:
return_type function_name(parameter_list) {
// body
}
Example:
int add(int a, int b) {
return a + b;
}
Types of Functions:
Library functions → printf(), scanf()
User-defined functions
Call by Value vs Call by Reference:
void swap(int *a, int *b) { int t=*a; *a=*b; *b=t; }
🔹 7. Arrays and Strings
Array:
int arr[5] = {1,2,3,4,5};
String:
char name[] = "Vini";
printf("%s", name);
Multidimensional Array:
int mat[3][3];
🔹 8. Pointers
Definition: Variable that stores address of another variable
Syntax:
int *ptr, a = 10;
ptr = &a;
printf("%d", *ptr);
Pointer Arithmetic:
ptr+1, ptr-1 (depends on data type size)
Applications:
Dynamic memory allocation, arrays, functions, strings
🔹 9. Structures and Unions
Structure:
struct Student {
int id;
char name[20];
};
Union:
Shares memory among members
union Data {
int i;
float f;
char c;
};
Difference:
Structure Union
Separate memory for each member Shared memory
All members active Only one active
🔹 10. File Handling
File Operations:
fopen(), fclose(), fprintf(), fscanf(), fread(), fwrite()
Example:
FILE *fp = fopen("[Link]", "w");
fprintf(fp, "Hello File");
fclose(fp);
🔹 11. Dynamic Memory Allocation
Function Purpose
malloc() Allocate memory
calloc() Allocate & initialize to zero
realloc() Resize memory
Function Purpose
free() Deallocate memory
🔹 12. Key Concepts for Exams
Recursion
Command-line arguments
Preprocessor directives (#define, #include)
Enumerations (enum)
Storage classes
🧠 Common C Exam Questions
1. Difference between call by value and reference
2. Explain pointers and their applications
3. What is the difference between structure and union?
4. Explain dynamic memory management functions.
5. Write a program to find factorial using recursion.
🧩 Part B: Object-Oriented Programming (OOP)
🔹 1. Overview
OOP = Object-Oriented Programming Paradigm
→ Based on objects and classes
Main Features:
1. Encapsulation
2. Abstraction
3. Inheritance
4. Polymorphism
🔹 2. Object and Class
Class: Blueprint of an object
Object: Instance of a class
class Student {
public:
string name;
void display() { cout << "Hello " << name; }
};
int main() {
Student s1;
[Link] = "Vini";
[Link]();
}
🔹 3. Encapsulation and Data Hiding
Encapsulation → Binding data + functions
Data Hiding → Restrict direct access using private/protected
class Bank {
private:
int balance;
public:
void setBalance(int b) { balance = b; }
int getBalance() { return balance; }
};
🔹 4. Inheritance
Allows new classes to derive properties of existing ones.
class A { public: void show() { cout << "A"; } };
class B: public A { public: void display() { cout << "B"; } };
Types:
Single
Multiple
Multilevel
Hierarchical
Hybrid
Access Specifiers: public, protected, private
🔹 5. Polymorphism
Compile-Time Polymorphism: Function & Operator Overloading
int add(int a, int b);
double add(double a, double b);
Runtime Polymorphism: Using Virtual Functions
class Base { public: virtual void show() { cout<<"Base"; } };
class Derived: public Base { public: void show() { cout<<"Derived"; } };
🔹 6. Abstraction
Shows only essential features, hides details.
→ Implemented via abstract classes or interfaces.
🔹 7. Constructors and Destructors
Constructor: Initialize objects automatically
class A {
public:
A() { cout << "Constructor"; }
~A() { cout << "Destructor"; }
};
🔹 8. Operator Overloading
Redefining operator meaning for user-defined types
class Complex {
int r, i;
public:
Complex operator+(Complex c) {
Complex temp;
temp.r = r + c.r;
temp.i = i + c.i;
return temp;
}
};
🔹 9. Templates (Generic Programming)
Allows writing type-independent code.
template <typename T>
T add(T a, T b) { return a + b; }
🔹 10. Exception Handling
try {
int x = 10/0;
} catch (...) {
cout << "Error occurred!";
}
🔹 11. File Handling in C++
ofstream fout("[Link]");
fout << "Hello";
[Link]();
ifstream fin("[Link]");
string line; fin >> line;
cout << line;
🧠 Common OOP Exam Questions
1. Explain the four main pillars of OOP.
2. What is the difference between compile-time and run-time polymorphism?
3. Write a C++ program to demonstrate inheritance.
4. How are constructors different from normal functions?
5. What is the use of templates in C++?
🔹 12. C vs C++ (OOP) Comparison
Feature C C++
Paradigm Procedural Object-Oriented
Data Security Less (global data) High (encapsulation)
Memory Management Manual Constructors/Destructors
Overloading Not supported Supported
Inheritance Not supported Supported
Feature C C++
Functions Top-down Bottom-up
🔹 13. Important PhD Entrance Focus Areas
1. C Programming Logic:
o Pointer arithmetic
o Memory allocation
o Recursion
o File operations
o Data structures (linked list, stack, queue basics)
2. OOP Conceptual Understanding:
o Class design principles
o Overloading vs overriding
o Virtual function behavior
o Multiple inheritance
o Templates and Exception Handling
3. Algorithmic Questions in C++:
o String reversal, sorting, recursion, factorial
o Object array manipulations
o Class and operator overloading programs
UNIT III Databases
Relational Databases, Query Language, E - R modeling, Normalization, Query Processing,
Transaction Processing, Integrity and Security;
1. Relational Databases (RDBMS)
🔹 Definition
A Relational Database stores data in tables (relations), consisting of rows (tuples) and
columns (attributes).
Developed by E. F. Codd (1970).
🔹 Basic Concepts
Term Meaning
Relation Table of data
Tuple Row / Record
Attribute Column / Field
Domain Set of allowed values for an attribute
Primary Key Uniquely identifies a tuple
Foreign Key Refers to a primary key in another table
Candidate Key All possible unique identifiers
Alternate Key Candidate key not chosen as primary key
🔹 Example
Student
RollNo (PK) Name
1 Vini
2 Lekshmi
🔹 Relational Algebra (Theoretical Query Language)
Operations:
Select (σ) → Select rows
σ(Marks > 80)(Student)
Project (π) → Select columns
π(Name, Dept)(Student)
Union (∪) → Combine tuples from two relations
Set Difference (-)
Cartesian Product (×)
Join (⋈)
Rename (ρ)
🔹 2. Query Language (SQL)
Structured Query Language (SQL) — standard language for RDBMS.
🔸 Data Definition Language (DDL)
CREATE TABLE, ALTER TABLE, DROP TABLE
CREATE TABLE Student (
RollNo INT PRIMARY KEY,
Name VARCHAR(50),
Dept VARCHAR(10),
Marks INT
);
🔸 Data Manipulation Language (DML)
INSERT, UPDATE, DELETE, SELECT
SELECT Name, Marks FROM Student WHERE Marks > 80;
🔸 Data Control Language (DCL)
GRANT, REVOKE
GRANT SELECT ON Student TO teacher;
🔸 Transaction Control Language (TCL)
COMMIT, ROLLBACK, SAVEPOINT
🔸 Joins
Type Description
INNER JOIN Common records from both tables
LEFT JOIN All from left + matching right
RIGHT JOIN All from right + matching left
FULL JOIN All from both sides
🔸 Subqueries
A query inside another query.
SELECT Name FROM Student WHERE Marks > (SELECT AVG(Marks) FROM Student);
🧠 3. E–R Modeling (Entity–Relationship Model)
Proposed by Peter Chen (1976) to represent real-world data as entities and relationships.
🔹 Components
Concept Symbol Description
Entity Rectangle Object with distinct existence (e.g., Student)
Attribute Oval Property of an entity (e.g., Name)
Relationship Diamond Association between entities
Primary Key Underlined Unique attribute
Weak Entity Double rectangle Depends on strong entity
🔹 Relationship Types
Type Example
One-to-One (1:1) Person–Passport
One-to-Many (1:N) Department–Students
Many-to-Many (M:N) Students–Courses
🔹 Cardinality and Participation
Defines how many instances of one entity relate to instances of another.
Total participation: every entity participates.
Partial participation: some entities participate.
🔹 Conversion of ER → Tables
1. Entity → Table
2. Attributes → Columns
3. Relationship → Foreign Key
4. M:N Relationship → Separate table
📐 4. Normalization
🔹 Purpose:
To remove data redundancy and anomalies (update, insert, delete).
🔹 Types of Anomalies
Insertion anomaly
Update anomaly
Deletion anomaly
🔹 Normal Forms
Form Condition Description
1NF Atomic values only No repeating groups
2NF 1NF + no partial dependency Non-key attribute depends on full PK
3NF 2NF + no transitive dependency Non-key not depend on another non-key
BCNF Every determinant is a candidate key Stronger than 3NF
4NF No multivalued dependency Removes multi-value redundancy
5NF No join dependency Fully decomposed
🔹 Example
Unnormalized:
(StudentID, Course1, Course2)
1NF:
Separate Course into rows
3NF:
Separate Student and Course into two tables, linked by StudentID
⚙️5. Query Processing
Steps from Query to Result:
1. Parsing and Translation → Check syntax, convert SQL → relational algebra
2. Optimization → Choose best execution plan (based on cost)
3. Evaluation → Execute optimized plan
Query Optimizer: Reduces response time & resource usage.
🔹 Query Optimization Techniques
Use of indexes
Join reordering
Predicate pushdown
Caching of subqueries
🔄 6. Transaction Processing
A transaction = single logical unit of work (one or more SQL statements).
🔹 ACID Properties
Property Description
Atomicity All or nothing execution
Consistency Database remains valid after transaction
Isolation Transactions do not interfere
Durability Changes persist after commit
🔹 States of Transaction
Active → Partially Committed → Committed / Failed / Aborted
🔹 Concurrency Control
Ensures correctness in multi-user environments.
Techniques:
Lock-based protocols (shared/exclusive)
Timestamp ordering
Two-phase locking (2PL)
Deadlock detection/prevention
🔹 Recovery Techniques
Commit Log
Checkpoints
Undo/Redo Logs
🔐 7. Database Integrity and Security
🔹 Data Integrity
Ensures accuracy and consistency of data.
Types:
1. Entity Integrity: Primary key ≠ NULL
2. Referential Integrity: Foreign key must exist in referenced table
3. Domain Integrity: Values within valid range or type
🔹 Security Mechanisms
Type Technique
Authentication Username/password
Authorization GRANT, REVOKE privileges
Encryption Encode data
Auditing Track access and modifications
Backup & Recovery Prevent data loss
🔹 SQL Security Example
GRANT SELECT, UPDATE ON Student TO user1;
REVOKE UPDATE ON Student FROM user1;
🔹 Common Threats
SQL Injection
Data leakage
Unauthorized access
Insider misuse
🧠 8. Summary Table
Concept Key Idea Exam Focus
Relational Model Data in tables with keys Primary/foreign key, relational algebra
SQL DDL, DML, DCL, TCL Joins, subqueries
E–R Model Entities and relationships Cardinality, weak entity
Normalization Reduce redundancy 1NF–BCNF
Query Processing SQL → Algebra → Execution Optimization
Transactions ACID properties Commit, rollback
Integrity Valid and consistent data Constraints
Security Protect database Authentication, authorization
🧾 9. Common PhD Entrance Questions
1. Define relation, tuple, and attribute with example.
2. Explain the difference between 3NF and BCNF.
3. What are the ACID properties of transactions?
4. Draw an ER diagram for “University Database.”
5. Explain query optimization and its importance.
6. How does SQL ensure integrity and security?
7. Differentiate between DDL, DML, and DCL.
8. What is normalization? Why is it needed?
Operating Systems - Process Management,
Scheduling, Deadlocks, Memory Management, File Systems
1. Introduction to Operating System
🔹 Definition
An Operating System (OS) is system software that manages hardware and software
resources and provides services to users and applications.
Examples: Windows, Linux, Unix, macOS, Android.
🔹 Functions of an OS
1. Process Management
2. Memory Management
3. File Management
4. I/O Management
5. Device Management
6. Security & Protection
7. Error Detection & Accounting
⚙️2. Process Management
🔹 What is a Process?
A process is a program in execution.
Components of a process:
Code (text)
Data
Stack
Heap
Program Counter (PC)
Process Control Block (PCB)
🔹 Process States
State Meaning
New Being created
State Meaning
Ready Waiting for CPU
Running Currently executing
Waiting/Blocked Waiting for I/O or event
Terminated Finished execution
📊 State Transition Diagram:
New → Ready → Running → Waiting → Ready → Terminated
🔹 Process Control Block (PCB)
Contains information about process:
PID (Process ID)
Program Counter
CPU Registers
Scheduling info
Memory info
I/O status
🔹 Process Operations
Creation: fork() in UNIX
Termination: exit()
Suspension / Resumption
🔹 Inter-Process Communication (IPC)
Mechanisms:
Shared Memory
Message Passing
Pipes
Sockets
Purpose: Data exchange between processes.
🕓 3. CPU Scheduling
CPU scheduling decides which process runs next on the CPU.
🔹 Scheduling Criteria
CPU Utilization
Throughput
Turnaround Time
Waiting Time
Response Time
Fairness
🔹 Scheduling Algorithms
Type Algorithm Description
Non-Preemptive FCFS (First Come First Serve) Executes in order of arrival
SJF (Shortest Job First) Shortest burst time first
Preemptive SRTF (Shortest Remaining Time First) Preemptive version of SJF
Each process gets fixed time
RR (Round Robin)
quantum
Priority Scheduling Based on priority value
Separate queues for different types of
Multilevel Queue
processes
Multilevel Feedback Processes move between queues
Queue dynamically
🔹 Example: Round Robin
If Time Quantum = 2
Process A (5), B (3), C (4)
→ A(2), B(2), C(2), A(2), B(1), C(2), A(1)
Advantages: Fairness
Disadvantages: Overhead due to context switching.
⚠️4. Deadlocks
🔹 Definition
A deadlock occurs when two or more processes are waiting indefinitely for resources held by
each other.
🔹 Necessary Conditions (Coffman Conditions)
1. Mutual Exclusion: Resource is non-shareable.
2. Hold and Wait: Process holds one resource while waiting for others.
3. No Preemption: Resource can’t be forcibly taken away.
4. Circular Wait: Chain of processes waiting for each other’s resource.
🔹 Deadlock Handling Methods
Method Technique
Prevention Break one of the four conditions
Avoidance Dynamically examine resource allocation (e.g., Banker’s Algorithm)
Detection Allow deadlock and detect using wait-for graph
Recovery Terminate or preempt processes
🔹 Banker’s Algorithm (Avoidance)
Based on resource allocation safety.
A state is safe if all processes can complete without deadlock.
🔹 Deadlock Recovery
Terminate one or more processes.
Preempt resources and roll back.
🧠 5. Memory Management
🔹 Purpose
Efficiently allocate and manage main memory among multiple processes.
🔹 Memory Hierarchy
Registers → Cache → Main Memory (RAM) → Secondary (Disk)
🔹 Contiguous Memory Allocation
Techniques:
1. Fixed Partitioning
2. Dynamic Partitioning
3. Relocation & Compaction
🔹 Non-Contiguous Allocation
(a) Paging
Memory divided into fixed-size blocks:
→ Pages (logical) and Frames (physical)
Each process has a page table mapping pages → frames.
Advantages: No external fragmentation
Disadvantages: Internal fragmentation possible
(b) Segmentation
Divides process memory into logical units (code, stack, data).
Each segment has base and limit.
Advantages: Logical view of memory
Disadvantages: External fragmentation
🔹 Virtual Memory
Allows processes to execute even if they don’t fit in physical memory completely.
Concepts:
Demand Paging
Page Fault
Page Replacement Algorithms
Common Algorithms:
Algorithm Description
FIFO Replace oldest page
LRU Replace least recently used
Optimal Replace page not used for longest time
Clock Circular buffer version of LRU
🔹 Thrashing
Too much paging activity (low CPU utilization).
→ Reduce by increasing memory or adjusting degree of multiprogramming.
💾 6. File Systems
🔹 Definition
A File System manages creation, deletion, reading, and writing of files, and maintains
directory structure.
🔹 File Concepts
File: Named collection of data.
File Attributes: Name, type, size, protection, location, timestamps.
File Operations: create, open, read, write, close, delete.
🔹 Directory Structure
Type Description
Single-level All files in one directory
Two-level Separate directory per user
Tree-structured Hierarchical directories
Acyclic graph Shared subdirectories
General graph Cycles allowed (with links)
🔹 File Allocation Methods
Type Description Pros Cons
Contiguous Files stored in continuous blocks Fast access Fragmentation
Linked Each file block points to next No fragmentation Slow random access
Indexed Uses index block to track Fast direct access Overhead of index block
🔹 Free Space Management
Bitmaps
Linked lists
Grouping
Counting
🔹 File Access Methods
Sequential Access: Data read sequentially
Direct (Random) Access: Access by block number
Indexed Access: Index allows fast lookup
🔹 File Protection and Security
Access Control Lists (ACLs)
Password Protection
Encryption
Backups
🧮 7. Summary Chart
Concept Core Idea Key Algorithm / Term
Process Program in execution PCB, States
Scheduling CPU allocation FCFS, SJF, RR, Priority
Deadlocks Resource waiting cycle Banker’s, Wait-for graph
Memory Allocation & replacement Paging, LRU, Segmentation
File System Data organization on disk FAT, Indexed, Directory structures
🧠 8. Common PhD Entrance Questions
1. Define process and thread.
2. Differentiate between preemptive and non-preemptive scheduling.
3. Explain the four conditions of deadlock.
4. How does the Banker’s algorithm avoid deadlocks?
5. Compare paging and segmentation.
6. Explain page replacement algorithms.
7. What is thrashing and how can it be prevented?
8. Differentiate between contiguous and non-contiguous file allocation.
9. What are the main functions of the operating system?
10. Explain ACID properties vs OS recovery. (Sometimes cross-topic question with DBMS!)
🧾 9. Mnemonics & Quick Revision
Concept Mnemonic
M–H–N–C → Mutual Exclusion, Hold & Wait, No Preemption, Circular
Deadlock Conditions
Wait
C–T–T–W–R → CPU Utilization, Throughput, Turnaround, Waiting,
Scheduling Goals
Response
Memory
F–L–O–C → FIFO, LRU, Optimal, Clock
Replacement
OS Functions P–M–F–I–D–S → Process, Memory, File, I/O, Device, Security
UNIT IV Networking
TCP/IP model, Layers, Functions and Protocols; Security - Cryptography, Symmetric Key
and Public Key Algorithms;
1. Introduction to Computer Networks
A computer network is an interconnection of computers and devices that can communicate
and share resources (data, applications, hardware).
🔹 Goals of Networking
Resource sharing
Data communication
Reliability
Scalability
Security
🔹 Types of Networks
Type Description Example
PAN Personal Area Network Bluetooth
LAN Local Area Network Office, school
MAN Metropolitan Area Network City-level
WAN Wide Area Network Internet
🖧 2. TCP/IP Model Overview
The TCP/IP Model is the foundation of the Internet.
Developed by DARPA, it defines how data is transmitted from one device to another over a
network.
🔹 TCP/IP Model Layers
Equivalent OSI
Layer Function Common Protocols
Layer(s)
Application, HTTP, HTTPS, FTP,
1. Application Provides user services
Presentation, Session DNS, SMTP, POP3, SSH
Reliable communication
2. Transport Transport TCP, UDP
between hosts
Logical addressing and
3. Internet Network IP, ICMP, ARP, RARP
routing
4. Network Access Physical transmission of Ethernet, Wi-Fi, PPP,
Data Link + Physical
(Link) frames MAC
🧠 3. Layer Functions and Protocols
🔹 1. Application Layer
User interface for network communication.
Handles application-specific protocols.
Protocol Function
HTTP/HTTPS Web access
FTP/SFTP File transfer
SMTP/POP3/IMAP Email services
DNS Domain name resolution
Telnet/SSH Remote login
🔹 2. Transport Layer
Ensures end-to-end delivery of data.
Provides segmentation, flow control, and error control.
Protocol Description
TCP (Transmission Control
Connection-oriented, reliable, uses ACK, retransmission
Protocol)
Connectionless, fast, no reliability, used for streaming &
UDP (User Datagram Protocol)
gaming
TCP Features:
3-way handshake (connection setup)
Flow control (Sliding Window)
Congestion control
Error detection
🔹 3. Internet Layer
Responsible for logical addressing, routing, and packet delivery.
Protocol Function
IP (Internet Protocol) Addressing and routing
ICMP Error reporting and diagnostics (e.g., Ping)
ARP Maps IP address to MAC address
Protocol Function
RARP Maps MAC to IP
IPv4 Address Format: 32 bits (e.g., [Link])
IPv6 Address Format: 128 bits (e.g., 2001:0db8:85a3::7334)
🔹 4. Network Access Layer
Responsible for physical transmission of data over medium.
Defines frame formats, MAC addressing, error detection.
Protocol Function
Ethernet (IEEE 802.3) Wired LAN
Wi-Fi (IEEE 802.11) Wireless LAN
PPP, SLIP Point-to-point links
CSMA/CD Collision detection (Ethernet)
🌍 4. TCP/IP Communication Process
1. Application creates message.
2. Transport layer (TCP/UDP) segments it.
3. Internet layer adds IP header (logical addressing).
4. Network Access layer frames and sends bits.
5. At receiver → reverse order (decapsulation).
📊 Data Unit at Each Layer:
Layer Data Unit
Application Message
Transport Segment / Datagram
Internet Packet
Layer Data Unit
Network Access Frame / Bits
🔒 5. Network Security and Cryptography
🔹 What is Cryptography?
Cryptography is the science of securing information by transforming it into an unreadable
format (encryption) and back into readable form (decryption).
🔹 Key Concepts
Term Description
Plaintext Original readable message
Ciphertext Encrypted message
Encryption Converting plaintext → ciphertext
Decryption Converting ciphertext → plaintext
Key Secret information used for encryption/decryption
🔹 Types of Cryptography
1. Symmetric Key Cryptography
2. Asymmetric (Public Key) Cryptography
🔐 6. Symmetric Key Cryptography
🔹 Concept
Same key used for encryption and decryption.
Very fast but key distribution is difficult.
🔹 Diagram
Sender ----[Encrypt with Key K]----> Ciphertext ----[Decrypt with Key K]----> Receiver
🔹 Common Algorithms
Algorithm Type Key Size Notes
DES (Data Encryption Standard) Block cipher 56 bits Outdated
3DES Triple DES 168 bits More secure
AES (Advanced Encryption Standard) Block cipher 128/192/256 bits Industry standard
RC4 Stream cipher Variable Used in SSL (older)
Blowfish / Twofish Block cipher Variable Faster alternative to AES
🔹 Advantages
Fast and efficient for large data
Less computationally expensive
🔹 Disadvantages
Key distribution problem
Not suitable for secure communication between unknown parties
🔑 7. Asymmetric (Public Key) Cryptography
🔹 Concept
Different keys for encryption and decryption.
o Public Key (PK): Shared openly
o Private Key (SK): Kept secret
Sender --[Encrypt with Receiver’s Public Key]--> Ciphertext --> Receiver --[Decrypt with
Private Key]--> Plaintext
🔹 Common Algorithms
Algorithm Function Key Size
RSA Encryption + Digital 1024–4096 bits
Algorithm Function Key Size
signatures
Diffie–Hellman Key exchange Variable
ECC (Elliptic Curve 256 bits (≈ RSA 3072
High security with small keys
Cryptography) bits)
DSA (Digital Signature Algorithm) Authentication –
🔹 Advantages
Solves key distribution issue
Provides confidentiality, authentication, integrity, and non-repudiation
🔹 Disadvantages
Slower than symmetric encryption
🔄 8. Hybrid Cryptography
Combines symmetric and asymmetric methods.
Example: In SSL/TLS →
o Asymmetric (RSA) used to share a session key
o Symmetric (AES) used for fast data transfer
🧾 9. Common Security Concepts
Concept Description
Hash Function One-way function that produces fixed-length output (e.g., SHA-256)
Digital Signature Ensures authenticity and integrity
Digital Certificate Issued by CA (Certificate Authority)
SSL/TLS Protocol for secure web communication
Firewall Filters unauthorized access
Concept Description
VPN Encrypts private network traffic over public network
⚔️10. Comparison Table
Feature Symmetric Asymmetric
Key Type Same key for both Different keys
Speed Fast Slow
Security Lower Higher
Key Distribution Difficult Easy
Examples AES, DES RSA, ECC
🧮 11. Common PhD Entrance Questions
1. Explain the layers of the TCP/IP model and their functions.
2. Differentiate between TCP and UDP.
3. What is the purpose of IP addressing?
4. Describe the working of ARP and ICMP.
5. Compare symmetric and asymmetric cryptography.
6. Explain AES and RSA algorithms.
7. What is a digital signature?
8. What are the major security threats in networking?
9. Explain SSL/TLS protocol.
10. Write differences between IPv4 and IPv6.
📘 12. Quick Revision Mnemonics
Concept Mnemonic
TCP/IP Layers A–T–I–N → Application, Transport, Internet, Network Access
Concept Mnemonic
OSI Layers A–P–S–T–N–D–P → All People Seem To Need Data Processing
Symmetric Examples AD 3B → AES, DES, 3DES, Blowfish, RC4
Asymmetric Examples RED → RSA, ECC, Diffie–Hellman
🧠 13. Summary Chart
Layer Function Protocols
Application User interface HTTP, FTP, DNS
Transport End-to-end reliability TCP, UDP
Internet Routing, addressing IP, ICMP, ARP
Network Access Framing, MAC Ethernet, Wi-Fi
Cryptography Type Examples
Symmetric Single key AES, DES
Asymmetric Public/Private keys RSA, ECC
Computer Architecture - Instruction Set Architectures,
Arithmetic Operations, Pipelines and Hazards, Caches
1. Introduction to Computer Architecture
Computer Architecture deals with the structure and behavior of a computer system as seen
by the programmer — including:
Instruction set design
Data representation
Processor design (CPU)
Memory hierarchy
I/O mechanisms
🧱 2. Instruction Set Architecture (ISA)
🔹 Definition
An Instruction Set Architecture (ISA) is the interface between hardware and software,
defining:
The instructions the CPU can execute
Registers, data types, addressing modes, and memory organization
🔹 ISA Components
1. Instruction Formats – Binary layout of instruction (opcode + operands)
2. Addressing Modes – How operands are accessed
3. Registers – CPU storage locations
4. Data Types – Integer, floating-point, character, etc.
5. Instruction Types – Data movement, arithmetic, control, logic, I/O
🔹 Instruction Format Example
Field Function
Opcode Operation (e.g., ADD, LOAD)
Source Operand(s) Input data
Destination Operand Output result
Example (in assembly):
ADD R1, R2, R3 → R1 = R2 + R3
🔹 Types of ISAs (by operand location)
Type Description Example
Stack-based Operands on a stack JVM bytecode
Accumulator-based One implicit operand (accumulator) Early 8080
Register-based Uses general-purpose registers ARM, MIPS
Type Description Example
Memory-based Operands in memory Intel x86
🔹 Addressing Modes
Mode Example Meaning
Immediate MOV R1, #5 Operand is constant
Register ADD R1, R2 Operand in register
Direct MOV R1, [2000] Operand at memory address
Indirect MOV R1, [R2] Memory address in register
Indexed MOV R1, [R2 + offset] Address = base + index
Relative JMP LABEL Address = PC + offset
➕ 3. Arithmetic Operations
🔹 Representation of Numbers
Unsigned binary → Only positive numbers
Signed magnitude → MSB = sign bit
1’s complement → Invert bits for negative numbers
2’s complement → Invert bits + 1 (most common)
Floating point → Represent real numbers (IEEE 754)
🔹 Binary Addition Rules
A B Sum Carry
0 0 0 0
0 1 1 0
1 0 1 0
A B Sum Carry
1 1 0 1
🔹 Binary Subtraction using 2’s Complement
Example:
7 – 5 = 7 + (2’s complement of 5)
7 → 0111
5 → 0101 → 1011 (2’s complement)
0111 + 1011 = 1 0010 → Result = 0010 = 2
🔹 Multiplication (Booth’s Algorithm)
Used for signed binary multiplication
Reduces number of partial products using bit-pair recoding
Rules:
| Bit Pair | Action |
|-----------|--------|
| 00 | No operation |
| 01 | Add multiplicand |
| 10 | Subtract multiplicand |
| 11 | No operation |
🔹 Floating Point Arithmetic
IEEE 754 Format (Single Precision – 32 bits):
| Sign (1 bit) | Exponent (8 bits) | Mantissa (23 bits) |
Value = (-1)^sign × ([Link]) × 2^(exponent – 127)
Example:
01000000101000000000000000000000
= +1.25 × 2^2 = 5.0
🔄 4. Pipelining
🔹 Concept
Pipelining allows overlapping execution of multiple instructions, increasing throughput
(instructions per unit time).
Like an assembly line — while one instruction is executed, another is decoded, another is
fetched.
🔹 5-Stage Instruction Pipeline (RISC)
Stage Function
IF Instruction Fetch
ID Instruction Decode & Register Fetch
EX Execute / ALU operation
MEM Memory Access
WB Write Back
🔹 Pipeline Performance
Speedup (S):
S=Time for non-pipelinedTime for pipelined=n×T(k+n−1)TS = \frac{\text{Time for non-
pipelined}}{\text{Time for pipelined}} = \frac{n \times T}{(k + n -
1)T}S=Time for pipelinedTime for non-pipelined=(k+n−1)Tn×T
where:
n = number of instructions
k = number of pipeline stages
Ideal Speedup = k (if no hazards)
🔹 Pipeline Hazards
Problems that cause pipeline stalls.
Type Description Example
Structural Hazard Hardware resource conflict Memory or register access clash
Data Hazard Dependency between instructions Instruction needs result not yet computed
Control Hazard Branch or jump instructions Pipeline flush due to wrong prediction
🔹 Data Hazard Types
Type Meaning Example
RAW (Read After Write) True dependency I1: R1=R2+R3, I2: R4=R1+R5
WAR (Write After Read) Anti-dependency I1: R2=R3+R4, I2: R3=R5+R6
WAW (Write After Write) Output dependency I1: R2=R3+R4, I2: R2=R5+R6
🔹 Techniques to Handle Hazards
Forwarding / Bypassing → Solve RAW hazards
Stalling / Pipeline Bubble → Delay instruction
Branch Prediction → Reduce control hazards
Out-of-order Execution → Dynamic scheduling (Tomasulo’s algorithm)
🧠 5. Cache Memory
🔹 Definition
Cache memory is a small, high-speed memory between CPU and main memory.
It stores frequently accessed instructions/data to reduce average memory access time.
🔹 Cache Performance Metrics
Average Memory Access Time (AMAT):
AMAT=Hit Time+(Miss Rate×Miss Penalty)AMAT = \text{Hit Time} + (\text{Miss Rate} \
times \text{Miss Penalty})AMAT=Hit Time+(Miss Rate×Miss Penalty)
🔹 Cache Mapping Techniques
Type Description Example
Direct Mapping Each block maps to one cache line Simple, fast
Fully Associative Any block can go to any cache line Expensive, complex
Set-Associative Cache divided into sets, each with few lines Balance between speed and cost
🔹 Cache Write Policies
Policy Description
Write-Through Update both cache and memory
Write-Back Update cache only, write to memory later
Write-Allocate Load block to cache on write miss
No-Write-Allocate Write directly to main memory on miss
🔹 Cache Miss Types (3 Cs)
1. Compulsory Miss – First access to data
2. Capacity Miss – Cache too small
3. Conflict Miss – Collision in direct-mapped cache
🔹 Multi-Level Cache
Modern CPUs use L1, L2, and L3 caches:
L1 – Closest, fastest, smallest
L2 – Intermediate
L3 – Shared across cores (largest, slower)
🧾 6. Summary Table
Concept Key Idea Example
ISA Interface between hardware and software MIPS, ARM, x86
Arithmetic Binary & floating-point ops 2’s complement, Booth’s
Pipelining Overlapping instruction execution 5-stage pipeline
Hazards Pipeline stalls RAW, WAR, WAW
Cache High-speed temporary memory Direct-mapped, set-associative
🧮 7. Common PhD Entrance Questions
1. What is an ISA? Compare RISC and CISC.
2. Explain different instruction formats with examples.
3. What are addressing modes in CPU design?
4. Explain Booth’s algorithm for multiplication.
5. What are pipeline hazards and how are they handled?
6. Define data hazards with examples.
7. Explain cache mapping techniques.
8. Derive formula for Average Memory Access Time.
9. Discuss write-through vs write-back policies.
10. Compare static and dynamic branch prediction.
🧠 8. Quick Revision Mnemonics
Concept Mnemonic
Pipeline Stages F–D–E–M–W → Fetch, Decode, Execute, Memory, Write-back
Hazards S–D–C → Structural, Data, Control
Data Hazards RAW–WAR–WAW
Cache Misses 3Cs → Compulsory, Capacity, Conflict
🧾 9. Important Formulas
1. Pipeline Speedup:
S=n×T(k+n−1)TS = \frac{n \times T}{(k + n - 1)T}S=(k+n−1)Tn×T
2. Average Memory Access Time (AMAT):
AMAT=Hit Time+(Miss Rate×Miss Penalty)AMAT = \text{Hit Time} + (\text{Miss
Rate} \times \text{Miss Penalty})AMAT=Hit Time+(Miss Rate×Miss Penalty)
3. CPU Execution Time:
CPU Time=IC×CPI×Clock Cycle Time\text{CPU Time} = \text{IC} \times \text{CPI} \
times \text{Clock Cycle Time}CPU Time=IC×CPI×Clock Cycle Time
4. Effective CPI (with stalls):
Effective CPI=Base CPI+(Stall Cycles per instruction)\text{Effective CPI} = \text{Base
CPI} + (\text{Stall Cycles per instruction})Effective CPI=Base CPI+
(Stall Cycles per instruction)
📘 10. Summary Mind Map
Computer Architecture
→ ISA → Instruction Formats, Addressing Modes
→ Arithmetic → 2’s Complement, Booth’s Algorithm, Floating Point
→ Pipeline → Stages, Hazards, Speedup
→ Cache → Mapping, Miss Types, Policies
UNIT V Software Engineering
Analysis, Design, Coding, Testing and Maintenance, Metrics, Object Oriented Analysis and
Design;
1. Introduction to Software Engineering
Software Engineering: Application of systematic, disciplined, and quantifiable
approaches to software development, operation, and maintenance.
Goal: To produce high-quality software that meets user requirements, delivered on time
and within budget.
Key Characteristics of Software:
Intangible, complex, and evolves over time.
Custom-built or generic.
Hard to measure in physical terms.
Software Crisis:
Challenges in cost estimation, schedule, maintenance, and quality due to increasing
complexity.
2. Software Development Life Cycle (SDLC)
A structured process that includes various stages of software development.
Phases:
1. Requirement Analysis
2. System Design
3. Implementation (Coding)
4. Testing
5. Deployment
6. Maintenance
Common SDLC Models:
Waterfall Model
Iterative Model
Spiral Model
V-Model
Agile Model (Scrum, XP)
3. Software Analysis
Focuses on understanding what the software must do.
a. Requirement Engineering
Steps:
1. Elicitation: Gathering requirements from users.
2. Analysis: Understanding and refining them.
3. Specification: Documenting requirements (SRS).
4. Validation: Ensuring correctness and completeness.
5. Management: Handling changes over time.
SRS (Software Requirements Specification):
Functional requirements (what system should do)
Non-functional requirements (performance, security, usability, etc.)
4. Software Design
Focuses on how to build the system to meet requirements.
a. Design Principles
Modularity, Abstraction, Cohesion, Coupling, Encapsulation, and Reusability.
b. Types of Design
1. High-Level Design (Architectural Design):
o Overall system structure, modules, and data flow.
2. Low-Level Design (Detailed Design):
o Logic of individual modules.
c. Design Tools
Data Flow Diagrams (DFD)
Entity Relationship Diagrams (ERD)
Structure Charts
Flowcharts
UML Diagrams (in OOAD)
5. Coding (Implementation)
Actual conversion of design into executable code.
Best Practices:
Follow coding standards and guidelines.
Use version control systems (Git).
Ensure readability and modularity.
Programming Languages:
C, C++, Java, Python, etc., depending on project scope.
6. Software Testing
Ensures that software is defect-free and meets requirements.
a. Levels of Testing
1. Unit Testing – testing individual modules.
2. Integration Testing – testing combined modules.
3. System Testing – entire system functionality.
4. Acceptance Testing – validation by client.
b. Types of Testing
Functional Testing
Non-functional Testing (Performance, Security)
Regression Testing
Black-box vs White-box Testing
c. Testing Techniques
Equivalence Partitioning
Boundary Value Analysis
Decision Table Testing
7. Software Maintenance
Modifying software after delivery to correct faults, improve performance, or adapt to a new
environment.
Types of Maintenance:
1. Corrective – fix bugs.
2. Adaptive – adjust to environment changes.
3. Perfective – enhance performance.
4. Preventive – prevent future issues.
Challenges: Cost, regression defects, documentation updates.
8. Software Metrics
Quantitative measures to assess software quality, productivity, and performance.
a. Product Metrics:
Size: LOC (Lines of Code), Function Points.
Complexity: Cyclomatic complexity.
Quality: Defect density, Reliability, Maintainability.
b. Process Metrics:
Effort, Schedule variance, Productivity.
c. Project Metrics:
Cost, Time, Resource utilization.
9. Object-Oriented Analysis and Design (OOAD)
a. Object-Oriented Concepts
Class, Object, Inheritance, Encapsulation, Abstraction, Polymorphism.
b. Object-Oriented Modeling
Use Case Diagram
Class Diagram
Sequence Diagram
Activity Diagram
State Chart Diagram
c. Object-Oriented Design Principles
SOLID Principles
o Single Responsibility Principle
o Open/Closed Principle
o Liskov Substitution
o Interface Segregation
o Dependency Inversion
d. Benefits
Reusability, Modularity, Extensibility, Maintainability.
10. Software Project Management
Focuses on planning, executing, and monitoring software projects.
Activities:
Estimation (COCOMO model)
Scheduling (PERT/CPM)
Risk Management
Resource Allocation
11. Software Quality Assurance (SQA)
Ensures that processes and products meet quality standards.
SQA Activities:
Reviews, audits, walkthroughs.
Verification & Validation (V&V).
Adherence to ISO 9001, CMMI, and IEEE standards.
12. Software Process Models Comparison
Model Advantages Disadvantages
Waterfall Simple, structured Rigid, no feedback loops
Spiral Risk management Complex, costly
Agile Flexible, user involvement Less documentation
Model Advantages Disadvantages
V-Model Verification at every stage Not good for long projects
13. Emerging Trends
Agile and DevOps Integration
Model-Driven Engineering
AI-assisted Software Development
Continuous Integration/Continuous Deployment (CI/CD)
🧮 Exam Preparation Tips
1. Understand SDLC and models clearly.
2. Memorize key differences (Waterfall vs Agile, Verification vs Validation).
3. Revise UML diagrams for OOAD.
4. Learn basic metrics and testing types.
5. Practice short definitions — e.g., “SRS,” “Cyclomatic complexity,” “Encapsulation.”
Web Technology - Scripting Languages, Client - Server Applications, Database
Connectivity;
1. Introduction to Web Technology
Definition:
Web Technology refers to the tools and techniques used to communicate between devices over
the World Wide Web (WWW).
It enables the creation of web applications, websites, and online services.
Key Components:
Web Client: The browser (Chrome, Firefox) used by users.
Web Server: Hosts web pages and applications (e.g., Apache, Nginx).
Protocol: HTTP / HTTPS.
Web Application: Software that runs on a web server and is accessed through a browser.
2. Web Architecture
Client–Server Model:
Client: Sends request (HTML pages, scripts).
Server: Processes request and returns response.
HTTP (Hypertext Transfer Protocol):
Stateless protocol for communication.
Request methods: GET, POST, PUT, DELETE, HEAD, OPTIONS.
Response codes:
o 1xx – Informational
o 2xx – Success (200 OK)
o 3xx – Redirection
o 4xx – Client Error (404 Not Found)
o 5xx – Server Error (500 Internal Server Error)
3. Web Development Tiers
1. Front-End (Client Side): HTML, CSS, JavaScript.
2. Back-End (Server Side): PHP, Python, [Link], Java, [Link].
3. Database Tier: MySQL, MongoDB, Oracle, PostgreSQL.
🧩 4. Scripting Languages
A. Client-Side Scripting
Executed on the user’s browser before sending data to the server.
Common Languages:
JavaScript (mainly used)
VBScript (legacy)
Key JavaScript Features:
Interactivity (validation, animations).
DOM Manipulation.
Event Handling (onClick, onSubmit).
ES6+ features: let, const, arrow functions, promises, async/await.
Example:
<script>
function validate() {
let name = [Link]("username").value;
if (name == "") {
alert("Please enter username");
return false;
}
return true;
}
</script>
B. Server-Side Scripting
Executed on the server to generate dynamic content before sending it to the client.
Common Languages:
PHP
Python (Flask/Django)
Java (JSP/Servlets)
[Link]
[Link] (C#)
Key Tasks:
Process form data.
Manage sessions and cookies.
Interact with databases.
Handle authentication.
Example (PHP):
<?php
$name = $_POST['username'];
echo "Welcome, " . $name;
?>
5. Client–Server Applications
A. Definition
Applications where the client requests services and the server provides them.
The client and server communicate using protocols (HTTP, TCP/IP).
B. Examples
Web Browsers ↔ Web Servers
Mobile Apps ↔ APIs
Email (SMTP, POP3, IMAP)
C. Architecture
1. Two-Tier Architecture:
o Client ↔ Database Server (e.g., desktop DB app).
2. Three-Tier Architecture:
o Client ↔ Application Server ↔ Database Server.
o Most modern web apps follow this (e.g., HTML + PHP + MySQL).
D. RESTful Architecture
REST (Representational State Transfer) uses stateless communication.
Uses HTTP methods (GET, POST, PUT, DELETE).
JSON and XML used for data exchange.
Example API Request:
GET /api/students/101
Response: { "name": "Vini", "course": "PhD" }
6. Database Connectivity
A. Overview
Database Connectivity allows web applications to store, retrieve, and manipulate data
dynamically.
B. Steps for Database Connectivity
1. Establish Connection to Database.
2. Execute SQL Queries.
3. Fetch Results.
4. Close the Connection.
C. PHP with MySQL Example
<?php
$conn = mysqli_connect("localhost", "root", "", "student_db");
if (!$conn) {
die("Connection failed: " . mysqli_connect_error());
}
$sql = "SELECT * FROM students";
$result = mysqli_query($conn, $sql);
while($row = mysqli_fetch_assoc($result)) {
echo $row["name"]."<br>";
}
mysqli_close($conn);
?>
D. Python (Flask) with SQLite Example
import sqlite3
conn = [Link]('[Link]')
cur = [Link]()
[Link]("SELECT * FROM students")
for row in [Link]():
print(row)
[Link]()
7. Database Design in Web Applications
Concepts:
E–R Modeling: Entities, Attributes, Relationships.
Normalization: Remove redundancy (1NF, 2NF, 3NF).
Primary / Foreign Keys: Maintain referential integrity.
Example Table:
student_id name dept
1 Vini CSE
2 Lekshmi ECE
8. Security in Web Applications
Common Security Practices:
HTTPS for data encryption.
Input validation to prevent SQL Injection.
Use prepared statements for DB queries.
Session management and secure cookies.
Authentication (OAuth, JWT).
Example of SQL Injection Prevention (PHP):
$stmt = $conn->prepare("SELECT * FROM users WHERE email=?");
$stmt->bind_param("s", $email);
$stmt->execute();
9. Frameworks & Tools
Purpose Technology Examples
Front-End HTML, CSS, JS React, Angular, Vue
Back-End Server-side Django, Flask, Laravel, [Link]
Database SQL / NoSQL MySQL, MongoDB
API Testing Tools Postman, Swagger
Server Deployment Apache, Nginx, XAMPP
10. Modern Web Concepts
AJAX (Asynchronous JavaScript and XML): Fetch data without reloading page.
JSON (JavaScript Object Notation): Lightweight data format.
WebSockets: Real-time, bidirectional communication.
Progressive Web Apps (PWA): Offline-capable, installable apps.
Microservices Architecture: Application broken into independent modules.
🧮 Key Exam Points Summary
Topic Focus Area
Client-side Scripting JavaScript syntax, DOM, events
Server-side Scripting PHP/Python/[Link] logic, form handling
Client–Server Model HTTP requests/responses, REST APIs
Database Connectivity SQL queries, CRUD operations
Security Input validation, SQL injection prevention
Topic Focus Area
Frameworks Django, Laravel, [Link] basics
✅ PhD Entrance Exam Preparation Tips
1. Revise architecture models (2-tier vs 3-tier).
2. Understand how HTTP requests flow in web apps.
3. Practice simple programs for DB connectivity (PHP/MySQL, Python/SQLite).
4. Remember scripting language differences (client vs server).
5. Revise security measures — frequent MCQ topic.
Cloud Computing - Virtualization; Big Data Analytics, NoSQL.
1. Introduction to Cloud Computing
Definition:
Cloud Computing is a model for delivering computing resources (servers, storage, databases,
networking, software, analytics, and intelligence) over the internet — on-demand, pay-per-use.
NIST Definition:
“Cloud computing is a model for enabling ubiquitous, convenient, on-demand network access to
a shared pool of configurable computing resources that can be rapidly provisioned and released
with minimal management effort.”
2. Key Characteristics
1. On-Demand Self-Service: Users can provision resources automatically.
2. Broad Network Access: Available over the Internet from any device.
3. Resource Pooling: Shared infrastructure serving multiple users (multi-tenancy).
4. Rapid Elasticity: Resources can be scaled up or down dynamically.
5. Measured Service: Pay-as-you-go billing model.
3. Cloud Service Models
Model Description Examples
AWS EC2, Google
IaaS (Infrastructure Provides virtualized hardware resources
Compute Engine, Microsoft
as a Service) such as VMs, networks, and storage.
Azure
Provides development platforms and tools
PaaS (Platform as a Google App Engine, AWS
for app creation without managing
Service) Elastic Beanstalk
infrastructure.
SaaS (Software as a Provides software applications over the Gmail, Salesforce,
Service) internet. Microsoft 365
4. Cloud Deployment Models
Model Features
Public Cloud Shared and open for public use (e.g., AWS, Google Cloud).
Private Cloud Dedicated infrastructure for a single organization.
Hybrid Cloud Combination of public and private clouds.
Community Cloud Shared by organizations with similar requirements.
5. Virtualization in Cloud Computing
a. Definition
Virtualization is the creation of a virtual (not physical) version of hardware resources — such
as servers, networks, or storage — using a hypervisor.
b. Types of Virtualization
1. Server Virtualization: Multiple virtual servers on one physical server.
2. Storage Virtualization: Pooling physical storage into a single logical unit.
3. Network Virtualization: Virtual networks isolated from physical ones.
4. Desktop Virtualization: Running desktop OS on virtual machines remotely.
5. Application Virtualization: Applications run in isolated environments.
c. Hypervisors
Type Description Examples
Type 1 (Bare Metal) Runs directly on hardware. VMware ESXi, Microsoft Hyper-V
Type 2 (Hosted) Runs on top of a host OS. VirtualBox, VMware Workstation
d. Benefits
Efficient utilization of hardware resources.
Isolation and security.
Scalability and flexibility.
Easy recovery and migration.
🧠 6. BIG DATA ANALYTICS
a. Definition
Big Data refers to extremely large and complex datasets that traditional databases cannot handle
efficiently.
b. Characteristics (5 V’s)
1. Volume: Massive data size (TB, PB).
2. Velocity: Rapid data generation (real-time streams).
3. Variety: Structured, semi-structured, unstructured data.
4. Veracity: Data accuracy and reliability.
5. Value: Insights derived from data.
c. Big Data Technologies
1. Hadoop Ecosystem
o HDFS (Hadoop Distributed File System): Data storage.
o MapReduce: Parallel data processing model.
o YARN: Resource management.
o Hive / Pig: Query and scripting tools.
2. Apache Spark
o In-memory processing for faster analytics.
o Supports batch + real-time (stream) data.
o APIs in Python, Java, Scala.
3. Kafka / Flume: Data ingestion and streaming.
4. HBase / Cassandra: NoSQL data storage systems.
d. Big Data Analytics Techniques
1. Descriptive Analytics: What happened (summary statistics, dashboards).
2. Predictive Analytics: What will happen (machine learning models).
3. Prescriptive Analytics: What should be done (optimization).
Applications:
Healthcare (disease prediction)
Finance (fraud detection)
Marketing (customer segmentation)
Smart cities (sensor analytics)
7. NoSQL DATABASES
a. Definition
NoSQL = “Not Only SQL” — a class of database systems designed to handle large volumes of
unstructured or semi-structured data, offering flexibility, scalability, and performance.
Key Features:
Schema-less design.
Horizontal scaling.
High availability and fault tolerance.
Supports large, distributed datasets.
b. Types of NoSQL Databases
Type Structure Examples Use Case
Redis,
Key–Value Stores Key-value pairs Caching, Session management
DynamoDB
JSON/XML/BSON MongoDB, Content management, user
Document Stores
documents CouchDB profiles
Column-Family
Column-oriented tables Cassandra, HBase Analytics, time-series data
Stores
Social networks,
Graph Databases Nodes & relationships Neo4j
recommendation engines
c. SQL vs NoSQL Comparison
Feature SQL (Relational) NoSQL (Non-Relational)
Schema Fixed Dynamic / Flexible
Scalability Vertical Horizontal
BASE (Basically Available, Soft state, Eventually
Transactions ACID
consistent)
Query Language SQL API-based / Query-specific
Data Type Structured Unstructured / Semi-structured
MySQL,
Examples MongoDB, Cassandra
PostgreSQL
d. CAP Theorem
In distributed databases, it is impossible to guarantee all three simultaneously:
C – Consistency
A – Availability
P – Partition tolerance
So, systems trade off between:
CP (Consistent + Partition Tolerant): e.g., HBase
AP (Available + Partition Tolerant): e.g., Cassandra
🔗 8. Integration of Cloud, Big Data, and NoSQL
Modern cloud systems (like AWS, Azure, GCP) provide integrated environments:
AWS: EC2 (compute), S3 (storage), Redshift (data warehouse), DynamoDB (NoSQL).
Azure: HDInsight (Hadoop), CosmosDB (NoSQL).
Google Cloud: BigQuery (analytics), Firestore (NoSQL).
Big Data platforms often run on Cloud using virtualized infrastructure for scalability.
🧩 9. Applications
Cloud-based data storage (Google Drive, Dropbox)
Real-time analytics (Netflix, YouTube recommendations)
IoT & Sensor Data Analysis
Healthcare data prediction systems
E-commerce personalization (Amazon, Flipkart)
⚙️10. Exam-Oriented Key Concepts Summary
Topic Key Focus Points
Cloud Computing Characteristics, Service Models, Deployment Models
Virtualization Definition, Types, Hypervisors, Benefits
Big Data 5V’s, Hadoop, MapReduce, Spark
Analytics Descriptive, Predictive, Prescriptive
NoSQL Types, CAP Theorem, SQL vs NoSQL
Integration Cloud + Big Data + NoSQL usage scenarios
🎯 PhD Entrance Preparation Tips
1. Revise definitions and differences (e.g., IaaS vs PaaS vs SaaS, SQL vs NoSQL).
2. Memorize 5V’s of Big Data and CAP Theorem trade-offs.
3. Understand virtualization concepts and hypervisor types.
4. Learn Hadoop architecture and NoSQL database examples.
5. Focus on real-world applications and advantages — often asked in conceptual
questions.