Do you remember that I mentioned several ways to obtain data last time? This time we will begin with importing Csv Files.
Open a new Code Cell in Colab
import pandas as pd
file = pd.read_csv('/content/files/test') # Fill in the file directory between the green single quotes hereWith that, we have successfully imported a csv file. A Csv file imported by Pandas roughly looks like this in the Colab terminal:
Notice this small tip: in Colab, most files, variables, and arrays do not need print(), like this:
fileRun this cell with the button on the left, and the result printed will look like the picture below:
(Notice that this picture shows only part of it)
Here is another way to import a file without attaching Drive:
There is no explanation here; you can just copy the answer:
from pathlib import path
import pandas as pd
path = Path("folder") / "subfolder" / "file.txt"
file = pd.read_csv('path')
This lets Pandas read a csv file without depending on the connection between Google Drive and Colab.
Last time we explained Numpy's Array. As its higher-level replacement, Pandas also has a similar data store: the Pandas DataFrame, abbreviated df. Built on the Numpy Array, it makes computation faster and is easier to use.
import pandas as pd
files = pd.read_csv('/content/ml-ds/test')
df = pd.DataFrame(files)
With that, a simple DF has been created.
Next, let's look at some basic uses of a Pandas DF
import pandas as pd
files = pd.read_csv('/content/ml-ds/test')
df = pd.DataFrame(files)
View the first five rows:
df.head()
View the last five rows:
df.tail()
View a specific column in a row:
df.loc[0,"sepal_length"]
Format: dataframe name.loc[row number, "column name"]
View one column across all rows:
df.loc[: , "sepal_length"]
View the content within a range of rows/columns:
df.loc[1:3, "sepal_length":"petal_length"]
iloc and loc do not have a significant difference in function. iloc only does not rely on the column index name when retrieving a column; it uses the column index number instead, for example:
df.iloc[0:2, 6:8]
These are some less common uses. Now let's introduce some common ones:
import pandas as pd
files = pd.read_csv('/content/ml-ds/test')
df = pd.DataFrame(files)
View all rows for one index name:
df["proximity"]
Filter by a condition:
df["proximity"] == "island"
Sort a Column by size:
df[df["ocean_proximity"] == "island"].sort_values("median_income", ascending=False).head(10)After the Pnadas indexing techniques above, let's explain some Pandas data-cleaning techniques:
import pandas as pd
files = pd.read_csv('/content/ml-ds/test')
df = pd.DataFrame(files)
Example of removing duplicates:
df.drop_duplicates(inplace=True)
Remove N/A-type data
df.dropna(subset=['critical_column'], inplace=True)
Fill in blank data (using median here):
median_value = df['numeric_col'].median()
df['numeric_col'].fillna(median_value, inplace=True)
Sum up an entire column:
df['Total'] = df.sum(axis=1, numeric_only=True)
Sum multiple Columns:
df['Total'] = df[['Column_A', 'Column_B']].sum(axis=1)
Or:
df['Total'] = df.iloc[:, 0:3].sum(axis=1)
Here, most basic uses of Pandas have basically been covered. This episode was updated after I finished updating Chapter 4, so I finally caught up. If anything is missing, please email me with advice at the bottom of the main site. Thank you.
That is basically all the content of this chapter. Recently I have mainly been studying deeper Deep learning, mostly about Computer Vision, such as Ultrlytics and OpenCV. Later I will also produce several episodes of Deep Learning content, not limited to ML frameworks. Here is a preview: after Linear regression&gradient descent comes Scikit- Learn, and after that we basically enter Model Building. Then there will be tutorials on matplotlib and Seaborn (mostly charting), followed by Transformers. After Transformers we enter the field of Deep Learning, and finally I will give a light lesson on small-parameter LLMs (SLMs) under the Transformers framework. You can look forward to it. See you next time~