Data Science & Machine Learning: post #4559 — TG.ME

Example: Surveying people standing outside the nearest shopping mall. Easy and inexpensive, but can introduce sampling bias.

🔹 11. Sampling Bias

Sampling bias occurs when the method used to select a sample systematically favors certain members.

Example: Surveying only customers who voluntarily contacted customer support — those customers may have unusually positive or negative experiences.

🔹 12. Representative Sample

A representative sample resembles the population in important characteristics.

Example: If population is 60% Group A and 40% Group B, a representative sample of 1,000 might have approximately 600 Group A and 400 Group B.

🔹 13. Sampling Error

Even a properly selected random sample won't usually produce exactly same results as entire population. Difference between sample estimate and true population value is called sampling error.

Example: True population average = ₹50,000, Sample average = ₹49,500. Sampling error generally decreases as sample size increases.

🔹 14. Larger Sample ≠ Always Better

A larger sample is not automatically a representative sample.

Example: Biased sample → 100,000 observations can still mislead, Representative sample → 1,000 observations can be better.



Quality of sampling matters, not just sample size.



🔹 15. Sampling in Machine Learning

Sampling is commonly used when working with large datasets.

Example: With 10 million records, you might sample a subset to explore data, test preprocessing code, develop visualizations, debug pipeline, perform preliminary analysis.

🔹 16. Train-Test Sampling

Machine Learning datasets are commonly divided into Full Dataset → Train and Test. Training set is used to learn patterns, test set is used to evaluate performance on unseen data. A validation set may also be used.

Example: The key idea is that evaluation data should provide reliable estimate of how model performs on new observations.

🔹 17. Sampling and Class Imbalance

Suppose fraud dataset contains 99,000 legitimate transactions and 1,000 fraudulent transactions = 1% fraud.

Example: Careless sampling could produce sample containing very few or no fraud cases. Techniques such as stratified sampling can help preserve representation.

🔹 18. Sampling Techniques Comparison

Simple Random = Randomly select individuals

Systematic = Select every kth observation

Stratified = Sample from each subgroup

Cluster = Select groups/clusters

Convenience = Select easily accessible individuals

🔹 19. Real-World Data Science Example

Population: 1,000,000 customers, Need sample: 20,000 customers.

Example: If churn rates differ significantly across Basic Plan, Premium Plan, Enterprise Plan, you could use stratified sampling and sample from each plan to ensure sample reflects structure of population.

🔹 20. Common Mistakes

Assuming every sample is representative

Example: A sample can be large but biased.

Confusing population and sample

Example: Population = Entire group, Sample = Subset

Confusing parameter and statistic

Example: Population → Parameter, Sample → Statistic

Thinking random sampling eliminates every type of error

Example: Random sampling can reduce selection bias, but sampling variability can still occur.
❤3
August 27, 2026 992 8