🔹 11. Sampling Bias ⭐
Sampling bias occurs when the method used to select a sample systematically favors certain members.
Example: Surveying only customers who voluntarily contacted customer support — those customers may have unusually positive or negative experiences.
🔹 12. Representative Sample
A representative sample resembles the population in important characteristics.
Example: If population is 60% Group A and 40% Group B, a representative sample of 1,000 might have approximately 600 Group A and 400 Group B.
🔹 13. Sampling Error
Even a properly selected random sample won't usually produce exactly same results as entire population. Difference between sample estimate and true population value is called sampling error.
Example: True population average = ₹50,000, Sample average = ₹49,500. Sampling error generally decreases as sample size increases.
🔹 14. Larger Sample ≠ Always Better
A larger sample is not automatically a representative sample.
Example: Biased sample → 100,000 observations can still mislead, Representative sample → 1,000 observations can be better.
Quality of sampling matters, not just sample size.
🔹 15. Sampling in Machine Learning ⭐
Sampling is commonly used when working with large datasets.
Example: With 10 million records, you might sample a subset to explore data, test preprocessing code, develop visualizations, debug pipeline, perform preliminary analysis.
🔹 16. Train-Test Sampling
Machine Learning datasets are commonly divided into Full Dataset → Train and Test. Training set is used to learn patterns, test set is used to evaluate performance on unseen data. A validation set may also be used.
Example: The key idea is that evaluation data should provide reliable estimate of how model performs on new observations.
🔹 17. Sampling and Class Imbalance
Suppose fraud dataset contains 99,000 legitimate transactions and 1,000 fraudulent transactions = 1% fraud.
Example: Careless sampling could produce sample containing very few or no fraud cases. Techniques such as stratified sampling can help preserve representation.
🔹 18. Sampling Techniques Comparison
Simple Random = Randomly select individuals
Systematic = Select every kth observation
Stratified = Sample from each subgroup
Cluster = Select groups/clusters
Convenience = Select easily accessible individuals
🔹 19. Real-World Data Science Example
Population: 1,000,000 customers, Need sample: 20,000 customers.
Example: If churn rates differ significantly across Basic Plan, Premium Plan, Enterprise Plan, you could use stratified sampling and sample from each plan to ensure sample reflects structure of population.
🔹 20. Common Mistakes
❌ Assuming every sample is representative
Example: A sample can be large but biased.
❌ Confusing population and sample
Example: Population = Entire group, Sample = Subset
❌ Confusing parameter and statistic
Example: Population → Parameter, Sample → Statistic
❌ Thinking random sampling eliminates every type of error
Example: Random sampling can reduce selection bias, but sampling variability can still occur.
