Skip to main content
Engineering LibreTexts

15.4: Exploratory Data Analysis

  • Page ID
    117623
  • \( \newcommand{\vecs}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)

    \( \newcommand{\vecd}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash {#1}}} \)

    \( \newcommand{\dsum}{\displaystyle\sum\limits} \)

    \( \newcommand{\dint}{\displaystyle\int\limits} \)

    \( \newcommand{\dlim}{\displaystyle\lim\limits} \)

    \( \newcommand{\id}{\mathrm{id}}\) \( \newcommand{\Span}{\mathrm{span}}\)

    ( \newcommand{\kernel}{\mathrm{null}\,}\) \( \newcommand{\range}{\mathrm{range}\,}\)

    \( \newcommand{\RealPart}{\mathrm{Re}}\) \( \newcommand{\ImaginaryPart}{\mathrm{Im}}\)

    \( \newcommand{\Argument}{\mathrm{Arg}}\) \( \newcommand{\norm}[1]{\| #1 \|}\)

    \( \newcommand{\inner}[2]{\langle #1, #2 \rangle}\)

    \( \newcommand{\Span}{\mathrm{span}}\)

    \( \newcommand{\id}{\mathrm{id}}\)

    \( \newcommand{\Span}{\mathrm{span}}\)

    \( \newcommand{\kernel}{\mathrm{null}\,}\)

    \( \newcommand{\range}{\mathrm{range}\,}\)

    \( \newcommand{\RealPart}{\mathrm{Re}}\)

    \( \newcommand{\ImaginaryPart}{\mathrm{Im}}\)

    \( \newcommand{\Argument}{\mathrm{Arg}}\)

    \( \newcommand{\norm}[1]{\| #1 \|}\)

    \( \newcommand{\inner}[2]{\langle #1, #2 \rangle}\)

    \( \newcommand{\Span}{\mathrm{span}}\) \( \newcommand{\AA}{\unicode[.8,0]{x212B}}\)

    \( \newcommand{\vectorA}[1]{\vec{#1}}      % arrow\)

    \( \newcommand{\vectorAt}[1]{\vec{\text{#1}}}      % arrow\)

    \( \newcommand{\vectorB}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)

    \( \newcommand{\vectorC}[1]{\textbf{#1}} \)

    \( \newcommand{\vectorD}[1]{\overrightarrow{#1}} \)

    \( \newcommand{\vectorDt}[1]{\overrightarrow{\text{#1}}} \)

    \( \newcommand{\vectE}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash{\mathbf {#1}}}} \)

    \( \newcommand{\vecs}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)

    \(\newcommand{\longvect}{\overrightarrow}\)

    \( \newcommand{\vecd}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash {#1}}} \)

    \(\newcommand{\avec}{\mathbf a}\) \(\newcommand{\bvec}{\mathbf b}\) \(\newcommand{\cvec}{\mathbf c}\) \(\newcommand{\dvec}{\mathbf d}\) \(\newcommand{\dtil}{\widetilde{\mathbf d}}\) \(\newcommand{\evec}{\mathbf e}\) \(\newcommand{\fvec}{\mathbf f}\) \(\newcommand{\nvec}{\mathbf n}\) \(\newcommand{\pvec}{\mathbf p}\) \(\newcommand{\qvec}{\mathbf q}\) \(\newcommand{\svec}{\mathbf s}\) \(\newcommand{\tvec}{\mathbf t}\) \(\newcommand{\uvec}{\mathbf u}\) \(\newcommand{\vvec}{\mathbf v}\) \(\newcommand{\wvec}{\mathbf w}\) \(\newcommand{\xvec}{\mathbf x}\) \(\newcommand{\yvec}{\mathbf y}\) \(\newcommand{\zvec}{\mathbf z}\) \(\newcommand{\rvec}{\mathbf r}\) \(\newcommand{\mvec}{\mathbf m}\) \(\newcommand{\zerovec}{\mathbf 0}\) \(\newcommand{\onevec}{\mathbf 1}\) \(\newcommand{\real}{\mathbb R}\) \(\newcommand{\twovec}[2]{\left[\begin{array}{r}#1 \\ #2 \end{array}\right]}\) \(\newcommand{\ctwovec}[2]{\left[\begin{array}{c}#1 \\ #2 \end{array}\right]}\) \(\newcommand{\threevec}[3]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \end{array}\right]}\) \(\newcommand{\cthreevec}[3]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \end{array}\right]}\) \(\newcommand{\fourvec}[4]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \\ #4 \end{array}\right]}\) \(\newcommand{\cfourvec}[4]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \\ #4 \end{array}\right]}\) \(\newcommand{\fivevec}[5]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \\ #4 \\ #5 \\ \end{array}\right]}\) \(\newcommand{\cfivevec}[5]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \\ #4 \\ #5 \\ \end{array}\right]}\) \(\newcommand{\mattwo}[4]{\left[\begin{array}{rr}#1 \amp #2 \\ #3 \amp #4 \\ \end{array}\right]}\) \(\newcommand{\laspan}[1]{\text{Span}\{#1\}}\) \(\newcommand{\bcal}{\cal B}\) \(\newcommand{\ccal}{\cal C}\) \(\newcommand{\scal}{\cal S}\) \(\newcommand{\wcal}{\cal W}\) \(\newcommand{\ecal}{\cal E}\) \(\newcommand{\coords}[2]{\left\{#1\right\}_{#2}}\) \(\newcommand{\gray}[1]{\color{gray}{#1}}\) \(\newcommand{\lgray}[1]{\color{lightgray}{#1}}\) \(\newcommand{\rank}{\operatorname{rank}}\) \(\newcommand{\row}{\text{Row}}\) \(\newcommand{\col}{\text{Col}}\) \(\renewcommand{\row}{\text{Row}}\) \(\newcommand{\nul}{\text{Nul}}\) \(\newcommand{\var}{\text{Var}}\) \(\newcommand{\corr}{\text{corr}}\) \(\newcommand{\len}[1]{\left|#1\right|}\) \(\newcommand{\bbar}{\overline{\bvec}}\) \(\newcommand{\bhat}{\widehat{\bvec}}\) \(\newcommand{\bperp}{\bvec^\perp}\) \(\newcommand{\xhat}{\widehat{\xvec}}\) \(\newcommand{\vhat}{\widehat{\vvec}}\) \(\newcommand{\uhat}{\widehat{\uvec}}\) \(\newcommand{\what}{\widehat{\wvec}}\) \(\newcommand{\Sighat}{\widehat{\Sigma}}\) \(\newcommand{\lt}{<}\) \(\newcommand{\gt}{>}\) \(\newcommand{\amp}{&}\) \(\definecolor{fillinmathshade}{gray}{0.9}\)
    Learning Objectives

    By the end of this section you should be able to

    • Describe exploratory data analysis.
    • Inspect DataFrame entries through appropriate indexing.
    • Use filtering and slicing to obtain a subset of a DataFrame.
    • Identify Null values in a DataFrame.
    • Remove or replace Null values in a DataFrame.

    Exploratory data analysis

    Exploratory Data Analysis (EDA) is the task of analyzing data to gain insights, identify patterns, and understand the underlying structure of the data. During EDA, data scientists visually and statistically examine data to uncover relationships, anomalies, and trends, and to generate hypotheses for further analysis. The main goal of EDA is to become familiar with the data and assess the quality of the data. Once data are understood and cleaned, data scientists may perform feature creation and hypothesis formation. A feature is an individual variable or attribute that is calculated from the raw data in the dataset.

    Data indexing can be used to select and access specific rows and columns. Data indexing is essential in examining a dataset. In Pandas, two types of indexing methods exist:

    • Label-based indexing using loc[]: loc[] allows you to access data in a DataFrame using row/column labels. Ex: df.loc[row_label, column_label] returns specific data at the intersection of row_label and column_label.
    • Integer-based indexing using iloc[]: iloc[] allows you to access data in a DataFrame using integer-based indexes. Integer indexes can be passed to retrieve specific data. Ex: df.iloc[row_index, column_index] returns specific data at the index row_index and column_index.
    Checkpoint: Indexing a DataFrame
    Concepts in Practice: DataFrame indexing

    Given the following code, respond to the questions below.

        import pandas as pd
    
        # Create sample data
        data = {
          "A": ["a", "b", "c", "d"],
          "B": [12, 20, 5, -10],
          "C": ["C", "C", "C", "C"]
        }
    
        df = pd.DataFrame(data)
    
    Concepts in Practice \(\PageIndex{1}\): DataFrame indexing

    What is the output of print(df.iloc[0, 0])?

    1. a
    2. b
    3. IndexError
    Answer

    1. a. The element in the first row and the first column is a.

    Concepts in Practice \(\PageIndex{2}\): DataFrame indexing

    What is the output of print(df.iloc[1, 1])?

    1. a
    2. b
    3. 20
    Answer

    c. The element in the second row and the second column is 20 .

    Concepts in Practice \(\PageIndex{3}\): DataFrame indexing

    What is the output of print(df.loc[2, 'A'])?

    1. c
    2. C
    3. b
    Answer

    a. The element at row label 2 and column label A is c.

    Data slicing and filtering

    Data slicing and filtering involve selecting specific subsets of data based on certain conditions or index/label ranges. Data slicing refers to selecting a subset of rows and/or columns from a DataFrame. Slicing can be performed using ranges, lists, or Boolean conditions.

    • Slicing using ranges: Ex: df.loc[start_row:end_row, start_column:end_column] selects rows and columns within the specified ranges.
    • Slicing using a list: Ex: df.loc[[label1, label2, ...], :] selects rows that are in the list [label1, label2, ...] and includes all columns since all columns are selected by the colon operator.
    • Slicing based on a condition: df[condition] selects only the rows that meet the given condition.

    Data filtering involves selecting rows or columns based on certain conditions. Ex: In the expression df[df['column_name'] > threshold], the DataFrame df is filtered using the selection operator ([]) and the condition(df['column_name'] > threshold) that is passed. All entries in the DataFrame df where the corresponding value in the DataFrame is True will be returned.

    Checkpoint: Indexing on a flight dataset
    Concepts in Practice: DataFrame slicing and filtering

    Given the following code, respond to the questions.

        import pandas as pd
    
        # Create sample data
        data = {
          "A": [1, 2, 3, 4],
          "B": [5, 6, 7, 8],
          "C": [9, 10, 11, 12]
        }
    
        df = pd.DataFrame(data)
    
    Concepts in Practice \(\PageIndex{4}\): DataFrame slicing and filtering

    Which of the following returns the first three rows of column A?

    1. df.loc[1:3, "A"]
    2. df.loc[0:2, "A"]
    3. df.loc[0:3, "A"]
    Answer

    b. df.loc[0:2, "A"] returns rows 0 to 2 (inclusive) of column A.

    Concepts in Practice \(\PageIndex{5}\): DataFrame slicing and filtering

    Which of the following returns the second row?

    1. df.iloc[2]
    2. df[2]
    3. df.loc[1]
    Answer

    c. loc[1] returns the row with label 1 , which corresponds to the second row of the DataFrame.

    Concepts in Practice \(\PageIndex{6}\): DataFrame slicing and filtering

    Which of the following returns the first column?

    1. df.loc[0, :]
    2. df.iloc[:, 0]
    3. df.loc["A"]
    Answer

    b. iloc[:, 0] selects all rows corresponding to the column index 0, which equals returning the first column.

    Concepts in Practice \(\PageIndex{7}\): DataFrame slicing and filtering

    Which of the following results in selecting the second and fourth rows of the DataFrame?

    1. df[df.loc[:, "A"] % 2 == 0]
    2. df[df[:, "A"] % 2 == 0]
    3. df[df.loc["A"] % 2 == 0]
    Answer

    a. The condition returns all rows where the value in the column with label A is divisible by 2.

    Handling missing data

    Missing values in a dataset can occur when data are not available or are not recorded properly. Identifying and removing missing values is an important step in data cleaning and preprocessing. A data scientist should consider ethical considerations throughout the EDA process, especially when handling missing data. They might consider answering questions such as "Why are the data missing?", "Whose data are missing?", and "Considering the missing data, is the dataset still a representative sample of the population under study?". The functions below are useful in understanding and analyzing missing data.

    • isnull(): The isnull() function can be used to identify Null entries in a DataFrame. The return value of the function is a Boolean DataFrame, with the same dimensions as the original DataFrame with True values where missing values exist.
    • dropna(): The dropna() function can be used to drop rows with Null values.
    • fillna(): The fillna() function can be used to replace Null values with a provided substitute value. Ex: df.fillna(df.mean()) replaces all Null values with the average value of the specific column.

    To define a Null value in a DataFrame, you can use the np.nan value from the NumPy library. Functions that aid in identifying and removing null entries are described in the table below the following code.

        import pandas as pd
        import numpy as np
    
        # Create sample data
        data = {
          "Column 1": ["A", "B", "C", "D", "E"],
          "Column 2": [np.NAN, 200, 500, 0, -10],
          "Column 3": [True, True, False, np.NaN, np.NaN]
        }
        
        df = pd.DataFrame(data)
    Column 1 Column 2 Column 3
    0 A NaN True
    1 B 200.0 True
    2 C 500.0 False
    3 D 0.0 NaN
    4 E -10.0 NaN
    Table 15.5
    Function Example Output Explanation
    isnull()
    df.isnull()
    
    Column 1 Column 2 Column 3
    0 False True False
    1 False False False
    2 False False False
    3 False False True
    4 False False True

    The df.isnull() function returns a Boolean array with Boolean values representing whether each entry is Null.

    fillna()
    df["Column 2"] =\
    df["Column 2"].fillna(df["Column 2"]
      .mean())
    Column 1 Column 2 Column 3
    0 A 172.5 True
    1 B 200.0 True
    2 C 500.0 False
    3 D 0.0 NaN
    4 E -10.0 NaN

    Null values in Column 2 are replaced with the mean of non-Null values in the column.

    dropna()
    # Applied after the run 
    # of the previous row
    df = df.dropna()
    Column 1 Column 2 Column 3
    0 A 172.5 True
    1 B 200.0 True
    2 C 500.0 False

    All rows containing a Null value are removed from the DataFrame.

    Table 15.6: Null identification and removal examples.
    Concepts in Practice \(\PageIndex{8}\): Missing value treatment

    Which of the following is used to check the DataFrame for Null values?

    1. isnan()
    2. isnull()
    3. isnone()
    Answer

    b. isnull() returns a Boolean DataFrame representing whether data entries are Null or not.

    Concepts in Practice \(\PageIndex{9}\): Missing value treatment

    Assuming that a DataFrame df is given, which of the following replaces Null values with zeros?

    1. df.fillna(0)
    2. df.replacena(0)
    3. df.fill(0)
    Answer

    a. fillna() replaces all Null values with the provided value passed as an argument.

    Concepts in Practice \(\PageIndex{10}\): Missing value treatment

    Assuming that a DataFrame df is given, what does the expression df.isnull().sum() do?

    1. Calculates sum of the non-Null values in each column
    2. Calculates the number of Null values in the DataFrame
    3. Calculates the number of Null values in each column
    Answer

    c. The function sum() is applied to each column separately and sums up values in columns. The result is the number of Null values in each column.

    Programming practice with Google

    Use the Google Colaboratory document below to practice EDA on a given dataset.

    Google Colaboratory document


    This page titled 15.4: Exploratory Data Analysis was last modified on Fri, 24 Jul 2026 08:47:35 GMT and is shared under a CC BY 4.0 license and was authored, remixed, and/or curated by OpenStax via source content that was edited to the style and standards of the LibreTexts platform.