14.3: Files in Different Locations and Working with CSV Files
- Page ID
- 117615
\( \newcommand{\vecs}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)
\( \newcommand{\vecd}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash {#1}}} \)
\( \newcommand{\dsum}{\displaystyle\sum\limits} \)
\( \newcommand{\dint}{\displaystyle\int\limits} \)
\( \newcommand{\dlim}{\displaystyle\lim\limits} \)
\( \newcommand{\id}{\mathrm{id}}\) \( \newcommand{\Span}{\mathrm{span}}\)
( \newcommand{\kernel}{\mathrm{null}\,}\) \( \newcommand{\range}{\mathrm{range}\,}\)
\( \newcommand{\RealPart}{\mathrm{Re}}\) \( \newcommand{\ImaginaryPart}{\mathrm{Im}}\)
\( \newcommand{\Argument}{\mathrm{Arg}}\) \( \newcommand{\norm}[1]{\| #1 \|}\)
\( \newcommand{\inner}[2]{\langle #1, #2 \rangle}\)
\( \newcommand{\Span}{\mathrm{span}}\)
\( \newcommand{\id}{\mathrm{id}}\)
\( \newcommand{\Span}{\mathrm{span}}\)
\( \newcommand{\kernel}{\mathrm{null}\,}\)
\( \newcommand{\range}{\mathrm{range}\,}\)
\( \newcommand{\RealPart}{\mathrm{Re}}\)
\( \newcommand{\ImaginaryPart}{\mathrm{Im}}\)
\( \newcommand{\Argument}{\mathrm{Arg}}\)
\( \newcommand{\norm}[1]{\| #1 \|}\)
\( \newcommand{\inner}[2]{\langle #1, #2 \rangle}\)
\( \newcommand{\Span}{\mathrm{span}}\) \( \newcommand{\AA}{\unicode[.8,0]{x212B}}\)
\( \newcommand{\vectorA}[1]{\vec{#1}} % arrow\)
\( \newcommand{\vectorAt}[1]{\vec{\text{#1}}} % arrow\)
\( \newcommand{\vectorB}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)
\( \newcommand{\vectorC}[1]{\textbf{#1}} \)
\( \newcommand{\vectorD}[1]{\overrightarrow{#1}} \)
\( \newcommand{\vectorDt}[1]{\overrightarrow{\text{#1}}} \)
\( \newcommand{\vectE}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash{\mathbf {#1}}}} \)
\( \newcommand{\vecs}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)
\(\newcommand{\longvect}{\overrightarrow}\)
\( \newcommand{\vecd}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash {#1}}} \)
\(\newcommand{\avec}{\mathbf a}\) \(\newcommand{\bvec}{\mathbf b}\) \(\newcommand{\cvec}{\mathbf c}\) \(\newcommand{\dvec}{\mathbf d}\) \(\newcommand{\dtil}{\widetilde{\mathbf d}}\) \(\newcommand{\evec}{\mathbf e}\) \(\newcommand{\fvec}{\mathbf f}\) \(\newcommand{\nvec}{\mathbf n}\) \(\newcommand{\pvec}{\mathbf p}\) \(\newcommand{\qvec}{\mathbf q}\) \(\newcommand{\svec}{\mathbf s}\) \(\newcommand{\tvec}{\mathbf t}\) \(\newcommand{\uvec}{\mathbf u}\) \(\newcommand{\vvec}{\mathbf v}\) \(\newcommand{\wvec}{\mathbf w}\) \(\newcommand{\xvec}{\mathbf x}\) \(\newcommand{\yvec}{\mathbf y}\) \(\newcommand{\zvec}{\mathbf z}\) \(\newcommand{\rvec}{\mathbf r}\) \(\newcommand{\mvec}{\mathbf m}\) \(\newcommand{\zerovec}{\mathbf 0}\) \(\newcommand{\onevec}{\mathbf 1}\) \(\newcommand{\real}{\mathbb R}\) \(\newcommand{\twovec}[2]{\left[\begin{array}{r}#1 \\ #2 \end{array}\right]}\) \(\newcommand{\ctwovec}[2]{\left[\begin{array}{c}#1 \\ #2 \end{array}\right]}\) \(\newcommand{\threevec}[3]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \end{array}\right]}\) \(\newcommand{\cthreevec}[3]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \end{array}\right]}\) \(\newcommand{\fourvec}[4]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \\ #4 \end{array}\right]}\) \(\newcommand{\cfourvec}[4]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \\ #4 \end{array}\right]}\) \(\newcommand{\fivevec}[5]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \\ #4 \\ #5 \\ \end{array}\right]}\) \(\newcommand{\cfivevec}[5]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \\ #4 \\ #5 \\ \end{array}\right]}\) \(\newcommand{\mattwo}[4]{\left[\begin{array}{rr}#1 \amp #2 \\ #3 \amp #4 \\ \end{array}\right]}\) \(\newcommand{\laspan}[1]{\text{Span}\{#1\}}\) \(\newcommand{\bcal}{\cal B}\) \(\newcommand{\ccal}{\cal C}\) \(\newcommand{\scal}{\cal S}\) \(\newcommand{\wcal}{\cal W}\) \(\newcommand{\ecal}{\cal E}\) \(\newcommand{\coords}[2]{\left\{#1\right\}_{#2}}\) \(\newcommand{\gray}[1]{\color{gray}{#1}}\) \(\newcommand{\lgray}[1]{\color{lightgray}{#1}}\) \(\newcommand{\rank}{\operatorname{rank}}\) \(\newcommand{\row}{\text{Row}}\) \(\newcommand{\col}{\text{Col}}\) \(\renewcommand{\row}{\text{Row}}\) \(\newcommand{\nul}{\text{Nul}}\) \(\newcommand{\var}{\text{Var}}\) \(\newcommand{\corr}{\text{corr}}\) \(\newcommand{\len}[1]{\left|#1\right|}\) \(\newcommand{\bbar}{\overline{\bvec}}\) \(\newcommand{\bhat}{\widehat{\bvec}}\) \(\newcommand{\bperp}{\bvec^\perp}\) \(\newcommand{\xhat}{\widehat{\xvec}}\) \(\newcommand{\vhat}{\widehat{\vvec}}\) \(\newcommand{\uhat}{\widehat{\uvec}}\) \(\newcommand{\what}{\widehat{\wvec}}\) \(\newcommand{\Sighat}{\widehat{\Sigma}}\) \(\newcommand{\lt}{<}\) \(\newcommand{\gt}{>}\) \(\newcommand{\amp}{&}\) \(\definecolor{fillinmathshade}{gray}{0.9}\)By the end of this section you should be able to
- Demonstrate how to access files within a file system.
- Demonstrate how to process a CSV file.
Opening a file at any location
When only the filename is used as the argument to the open() function, the file must be in the same folder as the Python file that is executing. Ex: For fileobj = open("file1.txt") in files.py to execute successfully, the file1.txt file should be in the same folder as files.py.
Often a programmer needs to open files from folders other than the one in which the Python file exists. A path uniquely identifies a folder location on a computer. The path can be used along with the filename to open a file in any folder location. Ex: To open a file named logfile.log located in /users/turtle/desktop the following can be used:
fileobj = open("/users/turtle/desktop/logfile.log")
| Operating System | File location |
|
|---|---|---|
| Mac |
|
|
| Linux |
|
|
| Windows |
|
|
For each question, assume that the Python file executing the open() function is not in the same folder as the out.txt file.
Each question indicates the location of out.txt, the type of computer, and the desired mode for opening the file. Choose which option is best for opening out.txt.
/users/turtle/files on a Mac for reading
fileobj = open("out.txt")fileobj = open("/users/turtle/files/out.txt")fileobj = open("/users/turtle/files/out.txt", 'w')
- Answer
-
b. The path followed by the filename enables the file to be opened correctly and in read mode by default.
c:\documents\ on a Windows computer for reading
fileobj = open("out.txt")fileobj = open("c:/documents/out.txt", 'a')fileobj = open("c:/documents/out.txt")
- Answer
-
c. The use of forward slashes / replacing the Windows standard backslashes \ enables the path to be read correctly. The path followed by the filename enables the file to be opened and when reading a file read mode is preferable.
/users/turtle/logs on a Linux computer for writing
fileobj = open("out.txt")fileobj = open("/users/turtle/logs/out.txt")fileobj = open("/users/turtle/logs/out.txt", 'w')
- Answer
-
c. The file is created or overwritten and changes can be written into the file.
c:\proj\assets on a Windows computer in append mode
fileobj = open("c:\\proj\\assets\\out.txt", 'a')fileobj = open("c:\proj\assets\out.txt", 'a')
- Answer
-
a. Since the backslashes in the Windows path appear in the string argument, which usually tells Python that this is part of an escape sequence, the backslashes must be ignored using an additional backslash \ character.
Working with CSV files
In Python, files are read from and written to as Unicode by default. Many common file formats use Unicode such as text files (.txt), Python code files (.py), and other code files (.c,.java).
Comma separated value (CSV, .csv) files are often used for storing tabular data. These files store cells of information as Unicode separated by commas. CSV files can be read using methods learned thus far, as seen in the example below.
\n characters and cells separated by commas.Raw text of the file:
Title, Author, Pages\n1984, George Orwell, 268\nJane Eyre, Charlotte Bronte, 532\nWalden, Henry David Thoreau, 156\nMoby Dick, Herman Melville, 538
"""Processing a CSV file."""
# Open the CSV file for reading
file_obj = open("books.csv")
# Rows are separated by newline \n characters, so readlines() can be used to read in all rows into a string list
csv_rows = file_obj.readlines()
list_csv = []
# Remove \n characters from each row and split by comma and save into a 2D structure
for row in csv_rows:
# Remove \n character
row = row.strip("\n")
# Split using commas
cells = row.split(",")
list_csv.append(cells)
# Print result
print(list_csv)
The code's output is:
[['Title', ' Author', ' Pages'], ['1984', ' George Orwell', ' 268'], ['Jane Eyre', ' Charlotte Bronte', ' 532'], ['Walden', ' Henry David Thoreau', ' 156'], ['Moby Dick', ' Herman Melville', ' 538']]
Why does readlines() work for reading the rows in a CSV file?
readlines()reads line by line using the newline\ncharacter.readlines()is not appropriate for reading a CSV file.readlines()automatically recognizes a CSV file and works accordingly.
- Answer
-
a. The newline
\ncharacter indicates where line breaks are in a file and is used by thereadlines()function to tell lines apart.
For the code in the example, what would be the output for the statement print(list_csv[1][2])?
532268Jane Eyre
- Answer
-
b. The cell at the 2nd row and 3rd column contains
'268'.
What is the output of the following code for the books.csv seen above?
file_obj = open("books.csv")
csv_read = file_obj.readline()
print(csv_read)
['Title, Author, Pages\n', '1984, George Orwell, 268\n', 'Jane Eyre, Charlotte Bronte, 532\n', 'Walden, Henry David Thoreau, 156\n', 'Moby Dick, Herman Melville, 538'][['Title', ' Author', ' Pages'], ['1984', ' George Orwell', ' 268'], ['Jane Eyre', ' Charlotte Bronte', ' 532'], ['Walden', ' Henry David Thoreau', ' 156'], ['Moby Dick', ' Herman Melville', ' 538']]Title, Author, Pages
- Answer
-
c. The first line of the books.csv is read into
csv_read. The first line is'Title, Author, Pages', which is the same as the first row.
Files such as Word documents (.docx) and PDF documents (.pdf), image formats such as Portable Network Graphics (PNG, .png) and Joint Photographic Experts Group (JPEG, .jpeg or .jpg) as well as many other file types are encoded differently.
Some types of non-Unicode files can be read using specialized libraries that support the reading and writing of different file types.
- PyPDF is a popular library that can be used to extract information from PDF files.
- BeautifulSoup can be used to extract information from XML and HTML files. XML and HTML files usually contain unicode with structure provided through the use of angled <> bracket tags.
- python-docx can be used to read and write DOCX files.
- Additionally, csv is a built-in library that can be used to extract information from CSV files.
The file fe.csv contains scores for a group of students on a final exam. Write a program to display the average score.
Interactive Code
- fe.csv
-
student id, score
123, 95
213, 92
111, 86
555, 97
621, 99
777, 100
312, 84
391, 88
398, 87
444, 95 - Answer
-
# Open the CSV file for reading
file_obj = open("fe.csv")# Rows are separated by newline \n characters, so readlines() can be used to read in all rows into a string list
csv_rows = file_obj.readlines()list_csv = []summation = 0.0for n in range(1, len(csv_rows)):
# Split using commas
cells = csv_rows[n].split(",")
summation += float(cells[1])
# Print result
print("Average score: ", summation/(len(csv_rows)-1))n
# Close the file
nfile_obj.close()


