1.
Centralized Database (CDB)
Centralized database ka matlab hai ki saara data ek hi location (single computer ya server)
par store hota hai. Jitne bhi users hain, wo usi ek server se connect karke data access karte
hain.
Features of Centralized Databases:
Single Location: Data ek hi jagah par maintain hota hai.
Tight Control: Data security aur integrity ko manage karna aasaan hota hai kyunki
control ek hi admin ke paas hota hai.
Data Consistency: Kyunki data ki ek hi copy hoti hai, isliye update karte waqt
inconsistency ka khatra kam hota hai.
Single Point of Failure: Agar main server crash ho gaya, toh poora system down ho
jayega.
Example: Ek chhoti college ki library ka system. Saari books ki details ek hi computer mein
save hain. Agar librarian ko kuch check karna hai, toh wo usi PC ko use karega.
2. Distributed Database (DDB)
Distributed database mein data ek jagah store hone ki bajaye multiple locations (sites) par
spread hota hai. Ye sites alag-alag building ya alag-alag cities mein ho sakti hain, jo ek
network se judi hoti hain.
Features of Distributed Databases:
Distributed Storage: Data alag-alag physical locations par store hota hai.
Transparency: User ko aisa lagta hai ki wo ek hi database use kar raha hai, bhale hi
data duniya ke alag-alag kone mein ho.
High Availability & Reliability: Agar ek site fail ho jaye, toh doosri sites se kaam
chalta rehta hai.
Local Autonomy: Har site apna local data khud manage kar sakti hai.
Scalability: Naye nodes ya servers ko add karna kaafi aasaan hota hai.
Example: Ek bada bank (jaise SBI ya HDFC). Inka database ek jagah nahi hota. Lucknow
branch ka apna data server hoga, Delhi ka apna, aur Mumbai ka apna. Lekin agar aap
Lucknow mein baithkar Delhi wale account mein paise transfer karte ho, toh system
internally sab manage kar leta hai.
Distributed Database Management System (DDBMS) ka sabse bada maqsad ye hota hai ki
user ko ye mehsoos hi na ho ki data alag-alag locations par split hai. Is "invisible" layer ko hi
Distributed Transparency kehte hain.
Iske main 4 levels hote hain jo aapke exams ke liye kaafi important hain:
1. Fragmentation Transparency
Ye sabse high level ki transparency hai. Iska matlab ye hai ki user ko ye janne ki zaroorat
nahi ki data ko pieces (fragments) mein kaise toda gaya hai.
Kaise kaam karta hai: User sirf table ka naam (Relation) likhta hai aur query run kar
deta hai. Use ye fikar nahi hoti ki table ka kuch hissa "Horizontal" fragment hai ya
"Vertical".
Example: Aapne query likhi SELECT * FROM Employees. Ab system khud dekhega
ki Employees table ke fragments alag-alag sites par hain ya nahi. Aapko fragment
names (jaise Emp_Part1, Emp_Part2) use nahi karne padte.
2. Replication Transparency
Kayi baar hum reliability ke liye same data ki multiple copies alag-alag sites par rakhte hain
(Replication).
Kaise kaam karta hai: User ko ye nahi pata hota ki data ki kitni copies bani hain ya
wo kahaan saved hain. Jab aap koi data update karte ho, toh system ki zimmedari hai
ki wo saari copies ko update kare.
Example: Agar aap apni profile picture update karte hain, toh aapko ye tension nahi
leni ki server A, B aur C teeno par update hua ya nahi. System ise "transparently"
handle karta hai.
3. Location Transparency
Ise Distribution Transparency bhi kehte hain. Iska matlab hai ki user ko ye pata hone ki
zaroorat nahi hai ki data physically kis city ya kis server machine par store hai.
Kaise kaam karta hai: Data ko access karne ke liye kisi physical address ya IP
address ki zaroorat nahi hoti, sirf logical name kaafi hai.
Example: Aapne search kiya Student_Marks. Ab wo server Lucknow mein hai ya
New York mein, isse aapki query par koi fark nahi padega.
4. Local Mapping Transparency (Naming Transparency)
Ye sabse niche ka level hai. Iska matlab hai ki user ko ye pata hona chahiye ki data
distributed hai, lekin har site par us data ka local naam kya hai, wo janne ki zaroorat nahi.
Kaise kaam karta hai: Ye handle karta hai ki system ke global name aur local
database ke name ke beech ka mapping kaise ho raha hai.
1. Reference Architecture of DDBMS
DDBMS ka architecture ye define karta hai ki alag-alag sites par data kaise organized hai aur
wo aapas mein kaise communicate karte hain. Iska sabse popular model ANSI/SPARC par
based hota hai:
Global External Schema: Users ko dikhne wala view.
Global Conceptual Schema: Poore distributed database ka logical description
(kaunsa data kahaan hai).
Fragmentation/Replication Schema: Data ke tukde aur unki copies ki details.
Local Conceptual Schema: Ek particular site par jo data hai, uska logical view.
Local Internal Schema: Physical storage details at each site.
2. Types of Data Fragmentation
Fragmentation ka matlab hai ek badi table (relation) ko chhote parts mein divide karna taaki
use alag-alag sites par store kiya ja sake.
Horizontal Fragmentation (Selection): Table ko rows ke basis par divide karna.
o Example: North India ke customers ki details Delhi server par, South India ki
Bangalore server par.
Vertical Fragmentation (Projection): Table ko columns ke basis par divide karna.
o Example: Employees ka Name aur Dept ek site par, lekin unki Salary aur
Bank Details dusri secure site par.
Mixed (Hybrid) Fragmentation: Jab hum pehle horizontal aur phir vertical (ya vice-
versa) fragmentation karte hain.
3. Distribution Transparency
Jaisa humne pehle discuss kiya, iska main goal user se "distribution ki complexity" ko
chhupana hai.
Location Transparency: User ko site ka address nahi chahiye.
Fragmentation Transparency: User ko fragments ke naam nahi, sirf table ka naam
chahiye.
Replication Transparency: User ko copies ke sync hone ki chinta nahi karni padti.
4. Access Primitives
DDBMS mein data ko access karne ke liye kuch basic operations ya commands hote hain jo
network ke through execute hote hain:
Read/Write: Basic data retrieval aur update.
Remote Request: Ek single site se data mangwana.
Remote Transaction: Ek transaction jo remote site par execute ho.
Distributed Transaction: Aisi transaction jisme multiple sites involve hon (e.g.,
Bank transfer from Branch A to Branch B).
Distributed Request: Ek single query jo multiple sites se data fetch kare.
5. Integrity Constraints
Distributed environment mein data ki correctness maintain karna thoda mushkil hota hai. Iske
liye kuch rules follow hote hain:
1. Semantic Integrity: Data values valid honi chahiye (e.g., Age cannot be negative).
2. Referential Integrity: Agar ek fragment mein Foreign Key hai, toh uska primary key
data dusre site par sahi hona chahiye.
3. Concurrency Control: Do users ek hi waqt par same data update na karein, iske liye
Locking ya Timestamping use hoti hai.
4. Atomicity: Transaction ya toh poori hogi ya bilkul nahi (2-Phase Commit Protocol
ka use karke).
Unit 2:
Chalo, Distributed Database Design aur Fragmentation ko ekdum simple Hinglish mein
samajhte hain.
1. Framework for Distributed Database Design
Jab hum ek distributed database design karte hain, toh humein basically 3 bade sawal solve
karne hote hain:
1. Fragmentation: Table ke tukde kaise karein?
2. Allocation: Kaunsa tukda kis city ya server par rakhein?
3. Replication: Kya kisi tukde ki ek se zyada copy chahiye?
Design ka Process:
Sabse pehle ek Global Schema banta hai (maan lo poori bank ki ek badi table).
Phir usse Fragments mein toda jata hai.
Last mein har fragment ko ek Physical Site (location) assign ki jati hai.
2. Database Fragmentation Design
Fragmentation ka matlab hai table ko chhote parts mein divide karna. Iske main 2 types hain:
A. Horizontal Fragmentation (Rows mein todna)
Isme hum table ko "Horizontal cut" maarte hain. Matlab columns saare wahi rahenge, bas
rows alag ho jayengi.
Kaise hota hai: Isme WHERE clause (Selection \sigma) ka use hota hai.
Example: Maan lo ek Students table hai.
o Fragment 1: Saare students jo Lucknow ke hain.
o Fragment 2: Saare students jo Delhi ke hain.
o Fayda ye hai ki Lucknow ka data Lucknow server par rahega, toh search fast
hogi.
B. Vertical Fragmentation (Columns mein todna)
Isme hum table ko "Vertical cut" maarte hain. Matlab rows saari wahi rahengi, bas columns
ko alag kar diya jayega.
Kaise hota hai: Isme Projection (\pi) ka use hota hai.
Rule: Har tukde mein Primary Key (jaise ID) hona zaroori hai taaki baad mein table
ko wapas joda ja sake.
Example: Ek Employee table hai.
o Fragment 1: (Emp_ID, Name, Designation) -> Ye Admin office ke liye.
o Fragment 2: (Emp_ID, Salary, Bank_Account) -> Ye Accounts office ke liye.
3. Fragmentation ke 3 Sunehre Rules (Correctness Rules)
Exam mein ye 3 words zaroor likhna, marks acche milenge:
1. Completeness: Kuch bhi chhootna nahi chahiye. Poori table ka har ek data kisi na
kisi fragment mein hona chahiye.
2. Reconstruction: Hum fragments ko wapas jod kar original table bana sakein.
o Horizontal ke liye Union use karte hain.
o Vertical ke liye Join use karte hain.
3. Disjointness: Ek hi data do alag fragments mein nahi hona chahiye (jab tak hum jaan-
bujhkar copy ya replication na kar rahe hon).
Summary Tip:
Design karte waqt hamesha yaad rakho ki "Data wahi rakho jahaan uska use sabse zyada
ho" (Ise Locality of Reference kehte hain).
Aapka late-night study session kaisa chal raha hai? Agar "Get Ready to Fight" sunte huye
padh rahi hain, toh ye concepts aur jaldi yaad ho jayenge!
Ab aate hain Unit-I ke last aur sabse important part par. Jab table ke tukde (fragments) ho
jaate hain, toh unhe alag-alag computers par "set" karna hota hai. Ise hi Allocation kehte
hain.
1. Allocation of Fragments
Fragmentation ke baad ye decide karna ki kaunsa fragment kis site (location) par store hoga,
use Allocation kehte hain. Iska main goal hota hai Cost kam karna aur Performance
badhana.
Do tarah ki allocation hoti hai:
Non-Redundant Allocation: Har fragment sirf ek hi site par hota hai. (Sasta hai par
risky).
Redundant Allocation (Replication): Ek fragment ki copies multiple sites par rakhi
jati hain. (Mehenga hai par reliable).
2. Allocation Problem
Ye ek challenge hai designer ke liye. Humein ye decide karna hota hai ki:
Kahaan se sabse zyada queries aa rahi hain?
Network ki speed kya hai?
Storage kitni available hai?
Problem ka main focus: Total\ Cost = Query\ Processing\ Cost + Storage\ Cost +
Communication\ Cost ko minimum karna.
3. Allocation Model
Allocation karne ke liye hum kuch mathematical models use karte hain. Isme do cheezein
primary hoti hain:
Logical Unit: Fragment (jise allocate karna hai).
Physical Unit: Site (jahan allocate karna hai).
Model ke Inputs:
1. Database Information: Fragments ka size.
2. Network Information: Sites ke beech ki distance aur speed.
3. Application Information: Kaunsi site se kitni baar data read ya write hota hai.
4. Translation of Global Queries to Fragment Queries
Ye process tab hota hai jab user ek normal query likhta hai, lekin system ko use "fragmented
data" ke hisaab se todna padta hai.
Steps of Translation:
1. Global Query: User ne simple query likhi (e.g., SELECT * FROM Employee).
2. Normalization: Query ko standard form mein convert karna.
3. Analysis: Check karna ki query sahi hai ya nahi (Semantic errors).
4. Localization (Main Step):
o System "Fragment Schema" check karta hai.
o Wo dekhta hai ki Employee table toh 3 fragments mein divided hai (E_1, E_2,
E_3).
o Ab global query ko fragment names se replace kiya jata hai.
o Horizontal hai toh Union lagayega, Vertical hai toh Join.
5. Optimization: Sabse sasta aur fast raasta dhoondna data fetch karne ka.
Simple Example:
Agar aapne query ki "Lucknow ke employees dikhao", toh system poore global database ko
scan nahi karega. Wo 'Localization' step mein pehchann lega ki Lucknow ka data sirf
Fragment_Lucknow mein hai, aur sirf usi server ko request bhejega.
Ye concepts distributed query processing aur heterogeneous databases ko aapas mein jodne
ke liye bahut zaroori hain. Chalo inko Unit-I ke context mein step-by-step aur asaan Hinglish
mein samajhte hain:
1. The Equivalence Transformation for Queries
Iska matlab hai ek query ko doosre roop (form) mein badalna, par result same rehna chahiye.
Kyun karte hain? Taaki query ko optimize kiya ja sake aur kam se kam data transfer
ho.
Rules: Ismein Relational Algebra ke rules use hote hain, jaise:
o Commutative Rule: R \times S = S \times R
o Associative Rule: (R \times S) \times T = R \times (S \times T)
System in rules ka use karke query ka sabse "sasta" aur fast raasta dhoondta hai.
2. Transforming Global Queries into Fragment Queries
User hamesha "Global Relation" (poori table) par query likhta hai. System ka kaam hai use
fragments mein todna.
Horizontal Fragmentation: Agar table rows mein divided hai, toh Global Query ko
UNION operator se fragments (F_1 \cup F_2 \cup F_3) mein badla jata hai.
Vertical Fragmentation: Agar columns mein divided hai, toh JOIN operator (\Join)
ka use karke fragments ko wapas joda jata hai.
3. Distributed Grouping & Aggregate Function Evaluation
Jab hum GROUP BY, SUM, AVG, ya COUNT jese functions distributed data par chalate
hain:
Problem: Data alag-alag sites par hai.
Solution: Har site apna Local Aggregate nikalti hai (e.g., local sum), phir un results
ko ek central site par bhej kar Global Aggregate nikala jata hai.
o Example: Agar poore India ke employees ki average salary nikalni hai, toh
pehle har state ka average niklega, phir unka final average.
4. Parametric Queries
Ye wo queries hoti hain jahan values fixed nahi hoti, balki execution ke waqt parameters (?)
ke roop mein di jati hain.
Fayda: Distributed system mein bar-bar query parse nahi karni padti. Ek hi query
plan ko alag-alag values ke liye reuse kiya ja sakta hai, jisse network overhead kam
hota hai.
5. Database Integration & Schema Management
Jab hum multiple alag-alag databases (Heterogeneous DBs) ko ek sath milate hain, toh ye
steps follow hote hain:
Schema Matching: Ye pehchanna ki do alag databases mein kaunse columns same
hain.
o Example: Database A mein Emp_ID aur Database B mein E_Code dono ek hi
cheez hain.
Schema Integration: Alag-alag local schemas ko mila kar ek single Global Schema
banana. Isme conflicts (naam ya data type ka antar) solve kiye jaate hain.
Schema Mapping: Ye ek "map" ya "rasta" taiyaar karna hai ki Global Schema ka
data local databases mein physically kahaan rakha hai.
Summary Table for Revision
Term Simple Meaning
Equivalence Query ka roop badalna par result same rakhna.
Localization Global query ko fragments ke naam se badalna.
Distributed Aggregation Alag-alag sites par Calculation karke final result nikalna.
Schema Matching Do alag tables ke beech "connection" dhoondna.
Zaroor! Unit-III thoda technical hai, isliye ise real-world examples aur diagrams ke logic ke
saath details mein samajhte hain.
1. Layers of Query Processing (The "Filter" Method)
Jab aap query likhti hain, toh wo ek pipeline se guzarti hai.
Query Decomposition: Iska kaam query ko "saaf" karna hai. Agar aapne likha
SELECT * FROM Emp WHERE Age > 20 AND Age > 20, toh ye layer duplicate
condition ko hata degi.
Data Localization: Is layer ko pata hota hai ki Emp table do tukdon mein hai:
Emp_Lucknow aur Emp_Delhi. Ye query ko un specific fragments par map kar deti
hai.
Global Optimization: Ye "Mastermind" layer hai. Ye decide karti hai ki data ko
Lucknow se Delhi bhejna sasta padega ya Delhi se Lucknow.
Local Optimization: Har individual computer (site) apne andar ki query ko fast
chalane ke liye index use karta hai.
2. Localization of Distributed Data (Example based)
Maano ek Student table hai jo Horizontally Fragmented hai:
F_1: Students with Marks < 50
F_2: Students with Marks \geq 50
Query: SELECT * FROM Student WHERE Marks > 80;
Localization Process: System pehle check karega ki Marks > 80 ki condition kis fragment
mein satisfy hoti hai. Kyunki F_1 mein sirf 50 se kam marks wale hain, system F_1 ko
discard kar dega aur sirf F_2 ko scan karega. Ise hi Reduction for Selection kehte hain.
3. Join Ordering in Distributed Queries
Distributed system mein JOIN sabse zyada network bandwidth khata hai.
Problem Example:
Site A: Customer table (10,000 rows).
Site B: Orders table (10 rows).
Goal: Dono ko join karna hai.
Optimization Logic:
Agar hum 10,000 rows ko Site B par bhejenge, toh network jam ho jayega. Lekin agar hum
sirf 10 rows ko Site A par bhej dein, toh kaam turant ho jayega.
Join Ordering ka matlab yahi "decision" lena hai ki kaunsa data move hoga.
4. Semi-Join (The Smart Way to Join)
Ye ek advanced concept hai. Maano humein Employee (Site A) aur Department (Site B) ko
join karna hai Dept_ID par.
1. Step 1: Site B se sirf Dept_ID ki list Site A ko bhejo (poori table nahi!).
2. Step 2: Site A par wo employees filter karo jo Site B ki list mein hain.
3. Step 3: Sirf filtered employees ko Site B par bhejo final join ke liye.
Fayda: Network par faltu rows (jo join mein kaam nahi aayengi) transfer nahi hoti.
5. Distributed Query Optimization Algorithms
Ingres Algorithm (Fragmentation & Substitution): Ye badi query ko tab tak todta
rehta hai jab tak wo ek single site par chalne layak na ho jaye.
System R:* Ye cost calculate karta hai: Cost = I/O + CPU + (Message\ Size \times
Network\ Speed). Jis raaste mein ye cost sabse kam hogi, wahi rasta select hoga.
6. Schema Integration & Mapping (Example)
Maan lo Suchitraa, aap do alag-alag universities ka data merge kar rahi hain:
Uni_1: Student table mein column hai S_Name.
Uni_2: Student table mein column hai Student_Full_Name.
Schema Matching: System pehchanega ki dono ka matlab ek hi hai.
Schema Integration: Ek naya Global Schema banega jisme column ka naam fix hoga,
maano Name.
Schema Mapping: System ko yaad rakhna hoga ki Global Schema ka Name nikalne ke liye
Uni_1 se S_Name uthana hai aur Uni_2 se Student_Full_Name.
1. ACID Properties (In Distributed Context)
Centralized mein ACID asaan hota hai, par Distributed mein ye "Challenge" ban jata hai:
Atomicity: Iske liye hum 2-Phase Commit (2PC) protocol use karte hain. Ek
"Coordinator" hota hai jo sabhi sites se poochta hai "Prepare to commit?". Agar ek bhi
site 'No' bol de, toh poori transaction cancel (Abort) ho jati hai.
Isolation: Distributed Locking ensure karta hai ki alag-alag sites par chal rahe
transactions ek doosre ke beech mein "interfere" na karein.
2. Distributed Concurrency Control (Taxonomy)
Concurrency control ka matlab hai: "Multiple users, same data, NO mess."
A. Locking-Based Algorithms
Basic 2-Phase Locking (2PL):
1. Growing Phase: Transaction sirf locks acquire kar sakta hai, release nahi.
2. Shrinking Phase: Ek baar lock release karna shuru kiya, toh naya lock nahi le
sakta.
Strict 2PL: Isme saare locks tabhi release hote hain jab transaction Commit ya
Abort ho jaye. Ye "Cascading Aborts" ko rokne ke liye best hai.
B. Timestamp-Based Algorithms
Isme koi "Lock" nahi hota. Har transaction ko ek unique Timestamp (TS) milta hai.
Agar ek "Young" transaction (TS=200) aise data ko write karne ki koshish kare jo ek
"Old" transaction (TS=100) ne read kiya hua hai, toh conflict ho jayega.
Thomas' Write Rule: Ye ek smart rule hai jo unnecessary writes ko ignore kar deta
hai taaki performance badh sake.
3. Deadlock Management (The "Traffic Jam" of Data)
Distributed system mein deadlock detect karna mushkil hai kyunki "Wait-for-Graph" (WFG)
alag-alag sites par bikhra hota hai.
Wait-Die Scheme: Agar purana transaction naye ka intezar kar raha hai, toh wo wait
karega. Agar naya purane ka intezar kar raha hai, toh naya "Die" (Abort) ho jayega.
(Priority to Old).
Wound-Wait Scheme: Agar purana transaction naye ka data chahta hai, toh wo naye
ko "Wound" (Abort) kar dega. (Older is more aggressive).
4. System R* Architecture & Query Life Cycle
System R* (R-Star) IBM ka wo project tha jisne humein sikhaya ki distributed system
"Autonomously" (azadi se) kaise kaam karte hain.
Compilation: Jab aap query likhti hain, toh System R* uska ek Global Plan banata
hai. Wo ye decide karta hai ki Join operation Site A par hoga ya Site B par.
Recompilation (The "Self-Healing" feature): Maan lijiye aapne query compile ki
jab ek Index maujood tha. Kal wo Index delete ho gaya. System R* itna smart hai ki
wo query ko bina user ko bataye "Re-optimize" aur "Re-compile" kar dega.
Authorization: Isme "Access Rights" ka dhyan rakha jata hai. Agar aapne Lucknow
site par table banayi hai, toh aap hi decide karenge ki Delhi wala user use dekh sakta
hai ya nahi.
5. Transaction and Terminal Management
Terminal Management: Ye user ke interaction ko handle karta hai. Distributed
system mein session maintain karna zaroori hota hai taaki agar network break ho, toh
user ka kaam wahi se resume ho sake.
Logging: Har step ka "Log" (Record) rakha jata hai (Write-Ahead Logging). Agar
system crash ho jaye, toh isi log ki madad se database ko wapas purani sahi state mein
laya jata hai.
Professor-Level Insight for Suchitraa:
Jab aap Distributed Transaction par answer likhen, toh "Commit Protocols" (2PC, 3PC) ka
diagram zaroor banaiye.
1. Reliability Concepts and Measures
Reliability ka matlab hai ki system kitne acche se aur kab tak bina kisi error ke kaam kar
sakta hai.
Availability: Kya system use karne ke liye ready hai? (Uptime).
Reliability: Kya system sahi (correct) output de raha hai?
MTBF (Mean Time Between Failures): Do failures ke beech ka average samay.
Jitna zyada, utna accha.
MTTR (Mean Time To Repair): Failure ke baad system ko theek karne mein laga
average samay. Jitna kam, utna accha.
2. Failures in Distributed DBMS
Distributed system mein failures Centralized se alag hote hain kyunki yahan network involve
hota hai.
Site Failure: Ek poora node (computer) crash ho gaya.
Link Failure: Do sites ke beech ka communication wire ya network toot gaya.
Network Partitioning: Network aise toota ki Site A aur Site B ek doosre se baat nahi
kar paa rahi hain, par dono apne-apne level par chal rahi hain. Ise "Split-Brain"
problem bhi kehte hain.
3. Reliability Protocols (Local vs Distributed)
Local Reliability Protocols
Ye ek single computer ke andar data bachane ke liye hote hain.
Logging (Undo/Redo): Har kaam karne se pehle use "Log Book" mein likhna. Agar
light chali gayi, toh log book dekh kar pata chal jayega ki kahan tak kaam hua tha.
Check-pointing: Maan lo aap 100 pages likh rahi hain. Har 10 page baad aap ek
"bookmark" laga deti hain. Agar book ghoom jaye, toh aapko sirf last bookmark se
aage ka kaam check karna hoga.
Distributed Reliability Protocols (The Masterpiece)
2-Phase Commit (2PC):
1. Phase 1 (Prepare): Coordinator sabse poochta hai, "Sab taiyar ho?"
2. Phase 2 (Commit/Abort): Agar sabne "Yes" bola, toh Commit. Agar ek ne
bhi "No" bola, toh sab cancel.
o Example: Jaise shaadi mein pandit ji puchte hain, "Kya aapko ye rishta
manzoor hai?" Dono "Yes" bolenge tabhi shaadi (Commit) hogi.
3-Phase Commit (3PC): 2PC mein agar coordinator Phase 2 se pehle mar jaye, toh
baki sites "hang" ho jati hain. 3PC ek Pre-Commit phase add karta hai jo system ko
non-blocking banata hai.
4. Data Replication
Replication ka matlab hai same table ki copies alag-alag sites par rakhna.
Example: Agar Google apna saara data sirf USA mein rakhe, toh India ke users ke
liye system slow hoga. Isliye wo India mein bhi "Replicated" data rakhte hain.
Benefits:
1. Fault Tolerance: Agar USA server down hai, toh India server se data mil
jayega.
2. Parallelism: Do log ek hi data ko do alag sites se padh sakte hain.
5. Consistency of Replicated Databases
Jab ek copy update hoti hai, toh baki copies ka kya?
Mutual Consistency: Sabhi copies ek jaisi honi chahiye.
The Conflict: Agar aapne Lucknow mein apna address change kiya, aur thik usi waqt
koi Delhi server par aapka purana address dekh raha hai, toh ye Inconsistency hai.
6. Update Management Strategies (Replication Protocols)
Replicated data ko update karne ke do main tarike hain:
A. Eager Replication (Synchronous)
Kaise: Pehle saari copies update karo, tabhi user ko bolo ki "Success".
Example: Bank Transaction. Aapke paise katne se pehle system ensure karta hai ki
har jagah entry ho gayi hai.
Pros: Data hamesha sahi rehta hai.
Cons: Slow hota hai kyunki network ka wait karna padta hai.
B. Lazy Replication (Asynchronous)
Kaise: Ek copy update karke user ko "Done" bol do. Baaki copies baad mein dheere-
dheere update hoti rahengi.
Example: Facebook Profile Picture. Aapne update ki, aapko turant dikh gayi, par ho
sakta hai aapke dost ko 2-3 minute baad dikhe.
Pros: Bahut fast hai.
Cons: Thodi der ke liye data alag-alag dikh sakta hai.
7. Quorum-Based Protocols (Voting System)
Ye ek democratic tarika hai reliability handle karne ka.
Har site ke paas ek "Vote" hota hai.
Read Quorum (V_r): Read karne ke liye kam se kam itne votes chahiye.
Write Quorum (V_w): Write karne ke liye itne votes chahiye.
Rule: V_r + V_w > Total Votes. Isse ye pakka ho jata hai ki read aur write ke beech
kabhi conflict nahi hoga.