Overview of Class 12
Data Handling using Pandas & Matplotlib
SECTION – A (MCQs)
1. What type of error is returned by the
following statement?
import pandas as pd
pd.Series([10, 20, 30], index=['x', 'y'])
(a) Value Error
(b) Syntax Error
(c) Name Error
(d) Logical Error
Answer: (a) Value Error
2. Which of the following statement is
correct regarding a DataFrame?
(a) DataFrame represents a single 1D
column in memory.
(b) Size of a DataFrame is mutable.
(c) We cannot change the column names of a DataFrame.
(d) DataFrame values are always immutable.
Answer: (b) Size of a DataFrame is mutable.
3. To display the first four rows of a
DataFrame object ‘DF’, you may write:
(a) DF.Head(4)
(b) DF.tail(4)
(c) DF.head(4)
(d) DF.Head()
Answer: (c) DF.head(4)
4. In Python Pandas, while performing
arithmetic operations between two DataFrames, unmatched indices or missing
values result in ______ by default.
(a) Zero
(b) None
(c) NaN
(d) False
Answer: (c) NaN
5. To access a specific value at the 2nd
row and 1st column of a DataFrame DF using integer-location, you can write:
(a) DF.iloc[2, 1]
(b) DF.loc[2, 1]
(c) DF.iat[1, 0]
(d) DF.at[2, 1]
Answer: (c) DF.iat[1, 0]
6. To delete a row from a DataFrame
permanently, you may use the ______ function/statement.
(a) remove()
(b) drop()
(c) clear()
(d) delete()
Answer: (b) drop()
7. To iterate over vertical subsets
(columns) of a DataFrame, ______ function may be used.
(a) iterrows()
(b) iteritems()
(c) itercols()
(d) loop()
Answer: (b) iteritems()
8. Which function is used to find the
reciprocal addition or reverse addition of a DataFrame with a scalar or another
DataFrame?
(a) radd()
(b) add()
(c) rev()
(d) reverse()
Answer: (a) radd()
9. Which of the following would give the
same output as DF1.mul(DF2) where DF1 and DF2 are DataFrames?
(a) DF1 * DF2
(b) multiply(DF1, DF2)
(c) DF1.multiply(DF2)
(d) Both (a) and (c)
Answer: (d) Both (a) and (c)
10. ______ divides the total distribution
into ten equal parts.
(a) Quartile
(b) Decile
(c) Percentile
(d) Median
Answer: (b) Decile
11. Which of the following is not a valid
mathematical/statistical descriptive function in Pandas?
(a) std()
(b) var()
(c) median()
(d) total()
Answer: (d) total()
12. Which submodule is commonly imported to
draw plots and charts in Python?
(a) matplotlib.pyplot
(b) pandas.plot
(c) numpy.chart
(d) csv.plot
Answer: (a) matplotlib.pyplot
13. Which of the following is not a valid
plotting function of pyplot?
(a) scatter()
(b) hist()
(c) boxplot()
(d) circle()
Answer: (d) circle()
14. To change the color of the bars in a
bar chart, which argument is used?
(a) barcolor
(b) color
(c) fillcolor
(d) c
Answer: (b) color
15. The command used to give a label to the
Y-axis of a graph is ______.
(a) plt.title()
(b) plt.xlabel()
(c) plt.ylabel()
(d) plt.axis()
Answer: (c) plt.ylabel()
16. Which argument in hist() is used to
specify the number of bins in a histogram?
(a) number
(b) bins
(c) bucket
(d) range
Answer: (b) bins
17. In a CSV file, data values are
typically separated by a:
(a) Semicolon (;)
(b) Colon (:)
(c) Comma (,)
(d) Pipe (|)
Answer: (c) Comma (,)
18. Which function is used to write a
Pandas DataFrame into a CSV file?
(a) to_csv()
(b) write_csv()
(c) save_csv()
(d) export_csv()
Answer: (a) to_csv()
19. Assertion(A): to_csv() function is used
to store DataFrame contents into a CSV file.
Reason(R): Pandas does not support file handling operations directly without
external libraries.
(a) Both A and R are true and R is the
correct explanation of A.
(b) Both A and R are true and R is not the correct explanation of A.
(c) A is true but R is false.
(d) A is false and R is true.
Answer: (c) A is true but R is false
20. Assertion(A): A Pandas DataFrame is a
2-dimensional labeled data structure.
Reason(R): Size of a DataFrame is immutable (cannot be changed once created).
(a) Both A and R are true and R is the
correct explanation of A.
(b) Both A and R are true and R is not the correct explanation of A.
(c) A is true but R is false.
(d) A is false and R is true.
Answer: (c) A is true but R is false
SECTION - B (Very Short
Answer Questions - 2 Marks Each)
21. Create a vertical bar graph for the
following data using matplotlib. Add suitable labels.
Products = ['Laptop', 'Mouse', 'Keyboard', 'Monitor']
Sales = [1200, 4500, 3200, 1500]
import matplotlib.pyplot as plt
Products = ['Laptop', 'Mouse', 'Keyboard', 'Monitor']
Sales = [1200, 4500, 3200, 1500]
plt.bar(Products, Sales, color='skyblue')
plt.xlabel('Products')
plt.ylabel('Sales')
plt.title('Product Sales Report')
plt.show()
22. Write Python statements to create a
DataFrame for the following data:
Item: ['Pen', 'Pencil', 'Eraser']
Price: [10, 5, 3]
Quantity: [50, 100, 200]
import pandas as pd
data = {
"Item": ['Pen', 'Pencil',
'Eraser'],
"Price": [10, 5, 3],
"Quantity": [50, 100, 200]
}
df = pd.DataFrame(data)
print(df)
23. Write code to create a 1D ndarray of
size 8 with all elements as 1, but the fourth element is 50.
import numpy as np
arr = np.ones(8)
arr[3] = 50
print(arr)
24. Create an ndarray with values ranging
from 10 to 50 each spaced with a difference of 10.
import numpy as np
x = np.arange(10, 51, 10)
print(x)
25. Write the output of the following
command:
import numpy as np
B = np.array([1, 2, 3, 4]) + 5
print(B)
Output:
[ 6 7
8 9]
26. What is Decile?
Answer: Deciles are descriptive
statistics that divide a ranked dataset into 10 equal parts, allowing us to
specify the value below which a given percentage of observations fall.
27. Write a program to create a Series
object with 4 random floating-point values and having custom index: ['w', 'x',
'y', 'z']
import pandas as pd
import numpy as np
s = pd.Series(np.random.rand(4), index=['w', 'x', 'y', 'z'])
print(s)
28. Mention at least four functions used
for finding descriptive statistics or summary in Pandas.
Answer: count(), sum(), mean(), std()
(Other options: min(), max(), median(), var())
29. Find the output of the following
arithmetic operations:
import numpy as np
b = np.array([[4, 2], [6, 8]])
print(b * 2)
print(b - 1)
Output:
[[ 8 4]
[12 16]]
[[3 1]
[5
7]]
30. Find the output for the following NumPy
array slicing:
import numpy as np
arr = np.array([[10, 20, 30], [40, 50, 60], [70, 80, 90]])
print(arr[:2, 1:])
Output:
[[20 30]
[50 60]]
SECTION - C (Short Answer Questions -4 Marks Each)
31. Given a set of marks scored by 10
students: [45, 60, 55, 70, 85, 90, 65, 50, 75, 80]
Write code to create a simple histogram, a horizontal histogram, and a
histogram with 5 specific bins.
import numpy as np
import matplotlib.pyplot as plt
marks = np.array([45, 60, 55, 70, 85, 90, 65, 50, 75, 80])
# 1. Simple histogram
plt.hist(marks)
plt.show()
# 2. Horizontal histogram
plt.hist(marks, orientation='horizontal')
plt.show()
# 3. Histogram with 5 bins
plt.hist(marks, bins=5)
plt.show()
32. Write code to draw a scatter plot or a
line graph where X-axis values are [2, 4, 6, 8] and Y-axis values are [5, 10,
15, 20] with red colored dashed lines and circular markers.
import matplotlib.pyplot as plt
x = [2, 4, 6, 8]
y = [5, 10, 15, 20]
plt.plot(x, y, color='red', linestyle='--', marker='o')
plt.title('Line Chart Example')
plt.xlabel('X-Axis')
plt.ylabel('Y-Axis')
plt.show()
33. Consider a DataFrame Employee: Modify
it by performing the following commands:
1. Add a new column 'Bonus' with values [500, 1000, 1500, 2000].
2. Add a new row for employee ID 'E5' with values ('E5', 'Amit', 45000, 'HR').
Employee['Bonus'] = [500, 1000, 1500,
2000]
Employee.loc['E5'] = ['E5', 'Amit', 45000, 'HR']
34. Write a program to read data from a CSV
file named StudentData.csv from path C:\Data\StudentData.csv and display its
first 5 rows.
import pandas as pd
df = pd.read_csv('C:\\Data\\StudentData.csv')
print(df.head())
35. (a) What is the difference between a
Bar Chart and a Histogram? Give two key differences.
Answer:
1. Nature of Data: A Bar chart is used to compare categorical data, whereas a
Histogram represents the frequency distribution of continuous numerical data.
2. Spacing: In a Bar chart, there are gaps between the bars, whereas in a
Histogram, the bars touch each other as the data is continuous.
OR
(b) Create a DataFrame from a dictionary
containing student details (RollNo, Name, Marks1, Marks2). Perform the
following:
1. Add a column 'Total' by adding Marks1 and Marks2.
2. Display the minimum marks obtained in Marks1.
import pandas as pd
data = {
'RollNo': [101, 102, 103],
'Name': ['Rahul', 'Priya', 'Karan'],
'Marks1': [80, 75, 90],
'Marks2': [85, 88, 78]
}
df = pd.DataFrame(data)
df['Total'] = df['Marks1'] + df['Marks2']
print('Minimum in Marks1:', df['Marks1'].min())
SECTION – D [Long Answer Questions - 5 Marks]
36. Write a complete Python program using
Matplotlib to plot multiple line charts on a single common plot with legends,
titles, and axis labels for three datasets:
Set1 = [10, 20, 30, 40]
Set2 = [15, 25, 10, 35]
Set3 = [5, 15, 25, 45]
import numpy as np
import matplotlib.pyplot as plt
set1 = [10, 20, 30, 40]
set2 = [15, 25, 10, 35]
set3 = [5, 15, 25, 45]
x = np.arange(4)
plt.plot(x, set1, color='blue', label='Dataset 1', marker='o')
plt.plot(x, set2, color='green', label='Dataset 2', marker='s')
plt.plot(x, set3, color='orange', label='Dataset 3', marker='^')
plt.title('Multiple Line Chart Comparison')
plt.xlabel('X-Axis Values')
plt.ylabel('Y-Axis Values')
plt.legend(loc='upper left')
plt.show()
37. Consider a DataFrame df containing
columns ['Department', 'EmployeeName', 'Salary']. Write Python statements for
the following operations:
1. To print the maximum salary in each Department.
2. To fill all NaN values in the Salary column with 30000.
3. To set the index of the DataFrame to EmployeeName.
4. To display the department-wise average salary.
5. To count the total number of employees in the 'IT' department.
# 1. Print maximum salary in each
department
print(df.groupby('Department')['Salary'].max())
# 2. Fill NaN values in Salary with 30000
df['Salary'].fillna(30000, inplace=True)
# 3. Set index to EmployeeName
df.set_index('EmployeeName', inplace=True)
# 4. Display department-wise average salary
print(df.groupby('Department')['Salary'].mean())
# 5. Count the number of employees in 'IT' department
print(df[df['Department'] == 'IT']['Salary'].count())
🎯 Class 11 IP Pandas Practice Quiz
Test your knowledge on Pandas Series and DataFrames!
Q1. What type of error is returned by: import pandas as pd?
pd.Series([10, 20, 30], index=['x', 'y'])
Q2. Which of the following statement is correct regarding a DataFrame?
Q3. To display the first four rows of a DataFrame object 'DF', you may write:
Q4. In Python Pandas, while performing arithmetic operations between two DataFrames, unmatched indices or missing values result in ______ by default.
Q5. To access a specific value at the 2nd row and 1st column of a DataFrame DF using integer-location, you can write:
Q6. To delete a row from a DataFrame permanently, you may use the ______ function/statement.
Q7. To iterate over vertical subsets (columns) of a DataFrame, ______ function may be used.
Q8. Which function is used to find the reciprocal addition or reverse addition of a DataFrame with a scalar or another DataFrame?
Q9. Which of the following would give the same output as DF1.mul(DF2) where DF1 and DF2 are DataFrames?
Q10. ______ divides the total distribution into ten equal parts.