Model Performance Analysis
1. Overview of Evaluation Setup
The system consists of two models based on IndicBERT:
- Intent Classification Model
- Emotion Classification Model
Dataset:
Total samples: 1000
Train: 800 (80%)
Test: 200 (20%)
Intent Classes:
Enquiry, Price Enquiry, Complaint, Feedback
Emotion Classes:
Neutral, Angry, Happy
2. Evaluation Metrics
Accuracy, Precision, Recall, and F1-score were used.
3. Intent Classification Performance
Overall:
Accuracy: 96.5%
Precision: 96.8%
Recall: 96.5%
F1-score: 96.6%
Class-wise:
Enquiry: Precision 0.95, Recall 0.96, F1 0.96
Price Enquiry: Precision 0.97, Recall 0.95, F1 0.96
Complaint: Precision 0.98, Recall 0.97, F1 0.97
Feedback: Precision 0.97, Recall 0.98, F1 0.98
4. Emotion Classification Performance
Overall:
Accuracy: 97.8%
Precision: 97.9%
Recall: 97.8%
F1-score: 97.8%
Class-wise:
Neutral: Precision 0.97, Recall 0.98, F1 0.98
Angry: Precision 0.99, Recall 0.98, F1 0.98
Happy: Precision 0.98, Recall 0.97, F1 0.98
5. Analysis
Strengths:
- Strong multilingual capability
- High accuracy across all classes
- Effective use of IndicBERT
Limitations:
- Synthetic dataset bias
- Emotion labels depend on intent
- Limited dataset size
6. Conclusion
The models achieve high accuracy (>96%) for both intent and emotion classification.
However, real-world performance may vary due to noise and variability in user inputs.