Business statistics - introduction, sampling and sample designs, collection of primary and secondary data, measures of central tendency, measures of dispersion, simple correlation, regression analysis, chi-square test, probability - Question Bank

1. In probability, 'independent events' are events where:
A) The occurrence of one event affects the probability of the other.
B) The occurrence of one event does not affect the probability of the other.
C) Both events cannot occur at the same time.
D) One event is a subset of the other.
2. What does a negatively skewed distribution look like?
A) The tail on the right side is longer or stretched.
B) The tail on the left side is longer or stretched.
C) It is perfectly symmetrical.
D) It has multiple peaks.
3. Which data collection method involves direct observation of phenomena as they occur?
A) Surveys
B) Interviews
C) Questionnaires
D) Observation
4. Systematic sampling involves selecting elements from a population at regular intervals after a random start. What is the interval called?
A) Sampling fraction
B) Sampling error
C) Sampling interval (k)
D) Sampling frame
5. What is the sample space in probability?
A) A single outcome of an experiment.
B) The set of all possible outcomes of an experiment.
C) The probability of an event occurring.
D) An event that cannot occur.
6. If events A and B are mutually exclusive, then P(A ∩ B) is:
A) P(A) * P(B)
B) 1
C) 0
D) P(A) + P(B)
7. Which of the following represents the probability of event A OR event B occurring, P(A U B)?
A) P(A) + P(B)
B) P(A) * P(B)
C) P(A) + P(B) - P(A ∩ B)
D) P(A) / P(B)
8. The Chi-Square test can be used to test for:
A) Goodness of fit
B) Independence of attributes
C) Both goodness of fit and independence of attributes
D) Correlation between variables
9. In multiple regression analysis, how many independent variables are typically involved?
A) One
B) Two
C) More than one
D) Zero
10. If the correlation coefficient is close to zero, it suggests:
A) A strong positive linear relationship.
B) A strong negative linear relationship.
C) A weak or non-existent linear relationship.
D) A perfect linear relationship.
11. A scatter plot is used to visually represent:
A) The central tendency of a dataset.
B) The dispersion of a dataset.
C) The relationship between two quantitative variables.
D) The probability distribution of a variable.
12. Which measure of dispersion is expressed in the same units as the original data?
A) Variance
B) Standard Deviation
C) Range
D) Coefficient of Variation
13. The Interquartile Range (IQR) is a measure of dispersion that represents:
A) The difference between the highest and lowest values.
B) The difference between the first and third quartiles.
C) The average deviation from the mean.
D) The most frequent value.
14. Which measure of central tendency is calculated by summing all values and dividing by the count?
A) Median
B) Mode
C) Mean
D) Range
15. The median is preferred over the mean when the data is:
A) Symmetrical
B) Normally distributed
C) Skewed or contains outliers
D) Continuous
16. A questionnaire is a tool used in which data collection method?
A) Observation
B) Interviews
C) Surveys
D) Focus Groups
17. Which method of data collection involves asking questions to respondents?
A) Observation
B) Experimentation
C) Survey
D) Content Analysis
18. Convenience sampling is a type of non-probability sampling where samples are selected based on:
A) Random selection.
B) Proportional representation.
C) Ease of access and availability.
D) Pre-defined quotas.
19. Which of the following is a non-probability sampling technique?
A) Simple Random Sampling
B) Systematic Sampling
C) Quota Sampling
D) Stratified Random Sampling
20. Cluster sampling is most useful when:
A) The population is small and easily accessible.
B) The population is geographically dispersed and the cost of data collection is high.
C) The population can be easily divided into homogeneous groups.
D) A high degree of precision is required.
21. In stratified sampling, the strata should be:
A) Heterogeneous within themselves but homogeneous between each other.
B) Homogeneous within themselves but heterogeneous between each other.
C) As diverse as possible.
D) As similar as possible.
22. Which sampling method involves dividing the population into homogeneous subgroups (strata) and then randomly sampling from each stratum?
A) Simple Random Sampling
B) Systematic Sampling
C) Stratified Random Sampling
D) Cluster Sampling
23. If an event is impossible, its probability is:
A) 0
B) 0.5
C) 1
D) Undefined
24. If an event is certain to occur, its probability is:
A) 0
B) 0.5
C) 1
D) Undetermined
25. The probability of an event occurring is always between:
A) 0 and 1 (inclusive)
B) -1 and 1 (inclusive)
C) 0 and 100 (inclusive)
D) 0.5 and 1 (inclusive)
26. Probability is a measure of:
A) Certainty that an event will not occur.
B) The likelihood or chance that an event will occur.
C) The average outcome of an experiment.
D) The total number of possible outcomes.
27. A key assumption for the Chi-Square test of independence is that:
A) The data must be normally distributed.
B) The sample size must be small.
C) The expected frequencies in each cell should not be too small (usually >5).
D) The variables must be continuous.
28. The Chi-Square (χ²) test is primarily used for:
A) Testing the difference between means of two groups.
B) Testing the relationship between two categorical variables.
C) Estimating population parameters.
D) Measuring the dispersion of continuous data.
29. The 'b' in the regression equation Y = a + bX represents:
A) The intercept.
B) The slope or the change in Y for a one-unit change in X.
C) The correlation coefficient.
D) The standard error of the estimate.
30. In simple linear regression, the equation Y = a + bX represents:
A) The correlation between X and Y.
B) The probability distribution of Y.
C) The regression line, where Y is the dependent variable and X is the independent variable.
D) The variance of X.
31. Regression analysis is used to:
A) Determine the central tendency of a single variable.
B) Describe the relationship between two or more variables and predict the value of one variable based on others.
C) Measure the spread of data around the mean.
D) Test hypotheses about population proportions.
32. A correlation coefficient of -0.8 indicates:
A) A weak positive linear relationship.
B) A perfect negative linear relationship.
C) A strong negative linear relationship.
D) No linear relationship.
33. A correlation coefficient of +1 indicates:
A) No linear relationship between the variables.
B) A perfect negative linear relationship.
C) A perfect positive linear relationship.
D) A weak positive linear relationship.
34. The correlation coefficient (r) ranges from:
A) 0 to 1
B) -1 to 0
C) -1 to 1
D) 0 to infinity
35. Simple correlation analysis is used to measure:
A) The cause-and-effect relationship between two variables.
B) The strength and direction of the linear relationship between two variables.
C) The average value of a single variable.
D) The probability of an event occurring.
36. A standard deviation of zero indicates:
A) High variability in the data.
B) No variability in the data; all values are the same.
C) The data is skewed.
D) The mean is not representative.
37. The standard deviation is:
A) The square of the variance.
B) The square root of the variance.
C) The difference between the highest and lowest values.
D) The most frequent value.
38. Which measure of dispersion is the average of the squared differences from the Mean?
A) Standard Deviation
B) Range
C) Variance
D) Interquartile Range
39. The range of a dataset is calculated as:
A) The sum of all values divided by the number of values.
B) The middle value when data is ordered.
C) The difference between the maximum and minimum values.
D) The square root of the variance.
40. Measures of dispersion are used to quantify:
A) The central point of the data.
B) The typical value in the dataset.
C) The spread or variability of the data.
D) The relationship between two variables.
41. Which measure of central tendency is suitable for qualitative data like 'color' or 'brand name'?
A) Mean
B) Median
C) Mode
D) Range
42. The mode is the measure of central tendency that represents:
A) The exact center of the data.
B) The most common value in a dataset.
C) The average of the first and last values.
D) The spread of the data.
43. The median of a dataset is:
A) The most frequently occurring value.
B) The average of all values.
C) The middle value when the data is ordered.
D) The difference between the highest and lowest values.
44. Which measure of central tendency is most affected by extreme values (outliers)?
A) Median
B) Mode
C) Mean
D) Geometric Mean
45. Published reports from government agencies, trade associations, and previous research studies are examples of:
A) Primary data
B) Experimental data
C) Secondary data
D) Survey data
46. Which type of data is collected directly from the source for the first time?
A) Secondary data
B) Tertiary data
C) Primary data
D) Internal data
47. What is the main disadvantage of census data collection compared to sampling?
A) It is less accurate.
B) It is more time-consuming and expensive.
C) It is difficult to analyze.
D) It is prone to sampling errors.
48. A sample design that ensures every member of the population has an equal and independent chance of being selected is known as:
A) Non-probability sampling
B) Purposive sampling
C) Random sampling
D) Quota sampling
49. Which of the following best describes sampling in statistics?
A) Collecting data from every single member of a population.
B) Selecting a subset of individuals from a larger population to make inferences about the population.
C) Ignoring data that does not fit a preconceived hypothesis.
D) Using only qualitative data for analysis.
50. What is the primary goal of Business Statistics?
A) To analyze past business performance only.
B) To make informed decisions by interpreting data.
C) To predict future stock market prices with certainty.
D) To manage daily operational tasks efficiently.