Data Handling using Pandas - II — NCERT Solutions
CBSE · Class 12 · Informatics Practices
NCERT Solutions for Data Handling using Pandas - II, CBSE Class 12 Informatics Practices: 28 textbook questions solved step by step. Covers Exercise.
Interactive on Super Tutor
Studying Data Handling using Pandas - II? Get the full interactive chapter.
Quizzes, flashcards, AI doubt-solver and a step-by-step study plan — built for NCERT solutions and more.
Free trial, no card needed.
The first 14 solutions are open to read. The other 14 are free with a Super Tutor account.
Exercise
1Write the statement to install the python connector to connect MySQL i.e. pymysql.Show solution
The statement to install the MySQL connector pymysql is:
pip install pymysql2Explain the difference between pivot() and pivot_table() function?Show solution
pivot() is used to reshape a DataFrame when the entries for the specified index and columns are unique. It rearranges data from rows into a new table.
pivot_table() works like pivot(), but it can handle duplicate entries by applying an aggregate function such as mean, sum, max, or min. The default aggregate function is mean.
So, the main difference is:
- pivot() does not work with duplicate values for the chosen index/columns.
- pivot_table() can reshape data even when duplicates are present, by aggregating them.
3What is sqlalchemy?Show solution
sqlalchemy is a library used to interact with the MySQL database by providing the required credentials. It also provides the create_engine() function to establish the connection.
4Can you sort a DataFrame with respect to multiple columns?Show solution
Yes, a DataFrame can be sorted with respect to multiple columns. In sort_values(), we pass a list of column names in the by parameter.
For example, if marks in Science are the same, then sorting can be done using Hindi as the next column:
dfUT3 = df[df.UT == 3]
print(dfUT3.sort_values(by=['Science', 'Hindi']))Here, the data is first sorted by Science and then by Hindi.
5What are missing values? What are the strategies to handle them?Show solution
A missing value is a value that is not present in a DataFrame. It is denoted by NaN.
The two main strategies to handle missing values are:
- Drop the object (row) having missing values.
- Fill/estimate the missing value using an appropriate value, such as 0, 1, previous value, next value, average, minimum, or maximum.
6Define the following terms: Median, Standard Deviation and variance.Show solution
- Median: the middle value of the data after arranging it in order. If there are an even number of values, it is the average of the two middle values.
- Standard Deviation: the measure of how much the values vary from the mean; it is the square root of variance.
- Variance: the average of squared differences from the mean.
7What do you understand by the term MODE? Name the function which is used to calculate it.Show solution
Mode is the value that appears the most number of times in a dataset. The function used to calculate it in Pandas is mode().
8Write the purpose of Data aggregation.Show solution
The purpose of data aggregation is to transform a dataset and produce a single numeric value from an array. It helps summarize data using functions like max(), min(), sum(), count(), std(), and var().
9Explain the concept of GROUP BY with help on an example.Show solution
GROUP BY is used to split data into groups based on some criteria, then apply a function to each group, and combine the results.
Example: if we group the marks DataFrame by Name, we can find the first entry, size, or sum for each student.
g1 = df.groupby('Name')
print(g1.first())
print(g1.size())This creates separate groups for Raman, Zuhaire, Ashravy and Mishti, and then operations can be performed group-wise.
10Write the steps required to read data from a MySQL database to a DataFrame.Show solution
The steps to read data from a MySQL database to a DataFrame are:
- Install and import the required libraries, such as
pymysqlandsqlalchemy. - Create a connection engine using
create_engine(). - Use one of the Pandas functions to read the table/query, such as:
pd.read_sql_query(query, engine)pd.read_sql_table(table_name, engine)pd.read_sql(sql, engine)
- Store the result in a DataFrame.
Example:
engine = create_engine('mysql+pymysql://root:password@localhost:3306/database_name')
df = pd.read_sql_query('SELECT * FROM INVENTORY', engine)11Explain the importance of reshaping of data with an example.Show solution
Reshaping data means changing the structure of a DataFrame so that it becomes suitable for analysis. Pandas provides pivot() and pivot_table() for this purpose.
Example: the sales data of stores can be reshaped so that Store becomes the index and Year becomes the columns. Then sales values are easier to compare across years.
This is important because reshaping can make data more readable, organized, and easier to analyze.
12Why estimation is an important concept in data analysis?Show solution
Estimation is important in data analysis because missing values represent a loss of information. If we replace them using a suitable estimate such as 0, previous value, next value, or average, we can still perform analysis and get a good approximation of the actual results. Without estimation, some data may have to be dropped, which reduces the size of the dataset.
13Assuming the given table: Product. Write the python code for the following:Show solution
The required task is to write Python code for the given Product table and related operations. The table can be stored as a DataFrame using a dictionary and pd.DataFrame().
For example:
import pandas as pd
data = {
'Item': ['TV', 'TV', 'TV', 'AC'],
'Company': ['LG', 'VIDEOCON', 'LG', 'SONY'],
'Rupees': [12000, 10000, 15000, 14000],
'USD': [700, 650, 800, 750]
}
df = pd.DataFrame(data)This creates the DataFrame for the given table.
13(a)To create the data frame for the above table.Show solution
To create the DataFrame for the given Product table:
import pandas as pd
data = {
'Item': ['TV', 'TV', 'TV', 'AC'],
'Company': ['LG', 'VIDEOCON', 'LG', 'SONY'],
'Rupees': [12000, 10000, 15000, 14000],
'USD': [700, 650, 800, 750]
}
df = pd.DataFrame(data)
print(df)Free with a Super Tutor account
Free with a Super Tutor account
Free with a Super Tutor account
Free with a Super Tutor account
Free with a Super Tutor account
Free with a Super Tutor account
Free with a Super Tutor account
Free with a Super Tutor account
Free with a Super Tutor account
Free with a Super Tutor account
Free with a Super Tutor account
Free with a Super Tutor account
Free with a Super Tutor account
Free with a Super Tutor account
14 more solved questions in Data Handling using Pandas - II
They are free with a Super Tutor account, along with practice quizzes and flashcards for this chapter. Free to start, no card needed.
Frequently Asked Questions
What are the important topics in Data Handling using Pandas - II for CBSE Class 12 Informatics Practices?
Are these NCERT Solutions for Data Handling using Pandas - II free?
How should I revise Data Handling using Pandas - II for the CBSE Class 12 board exam?
Sources & Official References
- NCERT Official — ncert.nic.in
- CBSE Academic — cbseacademic.nic.in
- CBSE Official — cbse.gov.in
- National Education Policy 2020 — education.gov.in
Content is aligned to the official syllabus. Refer to the board website for the latest curriculum.
More resources for Data Handling using Pandas - II
Practice Quiz
Test yourself with a quick quiz
Important Questions
Exam-style questions with answers
Revision Notes
Key points for last-minute revision
Formula Sheet
The chapter's formulas in one place
Chapter Summary
Understand the chapter at a glance
Concept Maps
See how topics connect
Study Plan
Step-by-step plan for this chapter
Flashcards
Quick-fire cards for active recall
Syllabus
What topics to cover
For serious students
Get the full Data Handling using Pandas - II chapter — start free.
Quizzes, flashcards, an AI doubt solver and a study plan for CBSE Class 12 Informatics Practices. Free to start, no card needed.