Skip to main content
Chapter 3 of 7
NCERT Solutions

Data Handling using Pandas - II — NCERT Solutions

CBSE · Class 12 · Informatics Practices

NCERT Solutions for Data Handling using Pandas - II, CBSE Class 12 Informatics Practices: 28 textbook questions solved step by step. Covers Exercise.

93 questions86 flashcards7 formulas & key relations5 concepts

Interactive on Super Tutor

Studying Data Handling using Pandas - II? Get the full interactive chapter.

Quizzes, flashcards, AI doubt-solver and a step-by-step study plan — built for NCERT solutions and more.

Free trial, no card needed.

28 Questions Solved · 1 Section

The first 14 solutions are open to read. The other 14 are free with a Super Tutor account.

Exercise

1Write the statement to install the python connector to connect MySQL i.e. pymysql.Show solution

The statement to install the MySQL connector pymysql is:

pip install pymysql
2Explain the difference between pivot() and pivot_table() function?Show solution

pivot() is used to reshape a DataFrame when the entries for the specified index and columns are unique. It rearranges data from rows into a new table.

pivot_table() works like pivot(), but it can handle duplicate entries by applying an aggregate function such as mean, sum, max, or min. The default aggregate function is mean.

So, the main difference is:

  • pivot() does not work with duplicate values for the chosen index/columns.
  • pivot_table() can reshape data even when duplicates are present, by aggregating them.
3What is sqlalchemy?Show solution

sqlalchemy is a library used to interact with the MySQL database by providing the required credentials. It also provides the create_engine() function to establish the connection.

4Can you sort a DataFrame with respect to multiple columns?Show solution

Yes, a DataFrame can be sorted with respect to multiple columns. In sort_values(), we pass a list of column names in the by parameter.

For example, if marks in Science are the same, then sorting can be done using Hindi as the next column:

dfUT3 = df[df.UT == 3]
print(dfUT3.sort_values(by=['Science', 'Hindi']))

Here, the data is first sorted by Science and then by Hindi.

5What are missing values? What are the strategies to handle them?Show solution

A missing value is a value that is not present in a DataFrame. It is denoted by NaN.

The two main strategies to handle missing values are:

  1. Drop the object (row) having missing values.
  2. Fill/estimate the missing value using an appropriate value, such as 0, 1, previous value, next value, average, minimum, or maximum.
6Define the following terms: Median, Standard Deviation and variance.Show solution
  • Median: the middle value of the data after arranging it in order. If there are an even number of values, it is the average of the two middle values.
  • Standard Deviation: the measure of how much the values vary from the mean; it is the square root of variance.
  • Variance: the average of squared differences from the mean.
7What do you understand by the term MODE? Name the function which is used to calculate it.Show solution

Mode is the value that appears the most number of times in a dataset. The function used to calculate it in Pandas is mode().

8Write the purpose of Data aggregation.Show solution

The purpose of data aggregation is to transform a dataset and produce a single numeric value from an array. It helps summarize data using functions like max(), min(), sum(), count(), std(), and var().

9Explain the concept of GROUP BY with help on an example.Show solution

GROUP BY is used to split data into groups based on some criteria, then apply a function to each group, and combine the results.

Example: if we group the marks DataFrame by Name, we can find the first entry, size, or sum for each student.

g1 = df.groupby('Name')
print(g1.first())
print(g1.size())

This creates separate groups for Raman, Zuhaire, Ashravy and Mishti, and then operations can be performed group-wise.

10Write the steps required to read data from a MySQL database to a DataFrame.Show solution

The steps to read data from a MySQL database to a DataFrame are:

  1. Install and import the required libraries, such as pymysql and sqlalchemy.
  2. Create a connection engine using create_engine().
  3. Use one of the Pandas functions to read the table/query, such as:
  • pd.read_sql_query(query, engine)
  • pd.read_sql_table(table_name, engine)
  • pd.read_sql(sql, engine)
  1. Store the result in a DataFrame.

Example:

engine = create_engine('mysql+pymysql://root:password@localhost:3306/database_name')
df = pd.read_sql_query('SELECT * FROM INVENTORY', engine)
11Explain the importance of reshaping of data with an example.Show solution

Reshaping data means changing the structure of a DataFrame so that it becomes suitable for analysis. Pandas provides pivot() and pivot_table() for this purpose.

Example: the sales data of stores can be reshaped so that Store becomes the index and Year becomes the columns. Then sales values are easier to compare across years.

This is important because reshaping can make data more readable, organized, and easier to analyze.

12Why estimation is an important concept in data analysis?Show solution

Estimation is important in data analysis because missing values represent a loss of information. If we replace them using a suitable estimate such as 0, previous value, next value, or average, we can still perform analysis and get a good approximation of the actual results. Without estimation, some data may have to be dropped, which reduces the size of the dataset.

13Assuming the given table: Product. Write the python code for the following:Show solution

The required task is to write Python code for the given Product table and related operations. The table can be stored as a DataFrame using a dictionary and pd.DataFrame().

For example:

import pandas as pd

data = {
    'Item': ['TV', 'TV', 'TV', 'AC'],
    'Company': ['LG', 'VIDEOCON', 'LG', 'SONY'],
    'Rupees': [12000, 10000, 15000, 14000],
    'USD': [700, 650, 800, 750]
}

df = pd.DataFrame(data)

This creates the DataFrame for the given table.

13(a)To create the data frame for the above table.Show solution

To create the DataFrame for the given Product table:

import pandas as pd

data = {
    'Item': ['TV', 'TV', 'TV', 'AC'],
    'Company': ['LG', 'VIDEOCON', 'LG', 'SONY'],
    'Rupees': [12000, 10000, 15000, 14000],
    'USD': [700, 650, 800, 750]
}

df = pd.DataFrame(data)
print(df)
13(b)To add the new rows in the data frame.

Free with a Super Tutor account

13(c)To display the maximum price of LG TV.

Free with a Super Tutor account

13(d)To display the Sum of all products.

Free with a Super Tutor account

13(e)To display the median of the USD of Sony products.

Free with a Super Tutor account

13(f)To sort the data according to the Rupees and transfer the data to MySQL.

Free with a Super Tutor account

13(g)To transfer the new dataframe into the MySQL with new values.

Free with a Super Tutor account

14Write the python statement for the following question on the basis of given dataset:

Free with a Super Tutor account

14(a)To create the above DataFrame.

Free with a Super Tutor account

14(b)To print the Degree and maximum marks in each stream.

Free with a Super Tutor account

14(c)To fill the NaN with 76.

Free with a Super Tutor account

14(d)To set the index to Name.

Free with a Super Tutor account

14(e)To display the name and degree wise average marks of each student.

Free with a Super Tutor account

14(f)To count the number of students in MBA.

Free with a Super Tutor account

14(g)To print the mode marks BCA.

Free with a Super Tutor account

14 more solved questions in Data Handling using Pandas - II

They are free with a Super Tutor account, along with practice quizzes and flashcards for this chapter. Free to start, no card needed.

Frequently Asked Questions

What are the important topics in Data Handling using Pandas - II for CBSE Class 12 Informatics Practices?
Key topics in Data Handling using Pandas - II include Descriptive Statistics, Data Aggregations, Sorting a DataFrame, GROUP BY Functions. Study these first, then practise questions on each for the CBSE Class 12 board exam.
Are these NCERT Solutions for Data Handling using Pandas - II free?
The first 14 of the 28 solutions on this page are open to read. The other 14 are free with a Super Tutor account — signing up is free and needs no card.
How should I revise Data Handling using Pandas - II for the CBSE Class 12 board exam?
Learn the core ideas first, then work through the 93 practice questions on Data Handling using Pandas - II. Revise definitions regularly and use flashcards for quick recall before the exam.

Sources & Official References

Content is aligned to the official syllabus. Refer to the board website for the latest curriculum.

For serious students

Get the full Data Handling using Pandas - II chapter — start free.

Quizzes, flashcards, an AI doubt solver and a study plan for CBSE Class 12 Informatics Practices. Free to start, no card needed.