Data analysis ka aadha kaam hai sahi rows aur columns nikaalna. Is lesson me hum wahi sales dataset use karenge:
Setup
from io import StringIO
import pandas as pd
csv_data = """order_id,date,city,category,product,quantity,price
1001,2026-01-05,Delhi,Electronics,Headphones,2,1500
1002,2026-01-06,Mumbai,Clothing,T-Shirt,3,499
1003,2026-01-06,Delhi,Clothing,Jeans,1,1299
1004,2026-01-07,Bangalore,Electronics,Mouse,4,650
1005,2026-01-08,Mumbai,Grocery,Rice 5kg,2,420
1006,2026-01-09,Pune,Electronics,Keyboard,1,
1007,2026-01-10,Delhi,Grocery,Tea 1kg,5,380
1008,2026-01-11,Bangalore,Clothing,Jacket,1,2499
1009,2026-01-12,Pune,Grocery,Oil 1L,3,160
1010,2026-01-12,Mumbai,Electronics,Charger,2,899"""
df = pd.read_csv(StringIO(csv_data), parse_dates=["date"])
print(df)
Output
order_id date city category product quantity price
0 1001 2026-01-05 Delhi Electronics Headphones 2 1500.0
1 1002 2026-01-06 Mumbai Clothing T-Shirt 3 499.0
2 1003 2026-01-06 Delhi Clothing Jeans 1 1299.0
3 1004 2026-01-07 Bangalore Electronics Mouse 4 650.0
4 1005 2026-01-08 Mumbai Grocery Rice 5kg 2 420.0
5 1006 2026-01-09 Pune Electronics Keyboard 1 NaN
6 1007 2026-01-10 Delhi Grocery Tea 1kg 5 380.0
7 1008 2026-01-11 Bangalore Clothing Jacket 1 2499.0
8 1009 2026-01-12 Pune Grocery Oil 1L 3 160.0
9 1010 2026-01-12 Mumbai Electronics Charger 2 899.0
loc vs iloc
|
| |
|---|---|---|
Kis se select karta hai | Label (index ka naam, column ka naam) | Position (0, 1, 2...) |
Slice ka end | Shamil hota hai | Shamil nahi hota |
Example |
|
|
Example
print(df.loc[2, "product"]) # index label 2, column "product"
print(df.iloc[2, 4]) # 3rd row, 5th column
print()
print(df.iloc[0:3, 2:5]) # pehli 3 rows, columns 2 se 4
print()
print(df.loc[0:2, ["city", "price"]]) # yaha 2 bhi shamil hai
Output
Jeans
Jeans
city category product
0 Delhi Electronics Headphones
1 Mumbai Clothing T-Shirt
2 Delhi Clothing Jeans
city price
0 Delhi 1500.0
1 Mumbai 499.0
2 Delhi 1299.0
Condition se filter
Condition ek True/False Series banati hai, aur usse DataFrame filter hota hai:
Example
delhi = df[df["city"] == "Delhi"]
print(delhi[["order_id", "product", "price"]])
Output
order_id product price
0 1001 Headphones 1500.0
2 1003 Jeans 1299.0
6 1007 Tea 1kg 380.0
Kai conditions ke liye & (and), | (or) use kariye — aur har condition ko brackets me rakhiye:
Example
costly_electronics = df[(df["category"] == "Electronics") & (df["price"] > 800)]
print(costly_electronics[["product", "price"]])
print()
mumbai_or_pune = df[(df["city"] == "Mumbai") | (df["city"] == "Pune")]
print(mumbai_or_pune[["city", "product"]])
Output
product price
0 Headphones 1500.0
9 Charger 899.0
city product
1 Mumbai T-Shirt
4 Mumbai Rice 5kg
5 Pune Keyboard
8 Pune Oil 1L
9 Mumbai Charger
isin(), between() aur query()
Example
print(df[df["city"].isin(["Delhi", "Bangalore"])][["city", "product"]])
print()
print(df[df["price"].between(400, 1000)][["product", "price"]])
print()
print(df.query("quantity >= 3 and city == 'Mumbai'")[["product", "quantity"]])
Output
city product
0 Delhi Headphones
2 Delhi Jeans
3 Bangalore Mouse
6 Delhi Tea 1kg
7 Bangalore Jacket
product price
1 T-Shirt 499.0
3 Mouse 650.0
4 Rice 5kg 420.0
9 Charger 899.0
product quantity
1 T-Shirt 3
loc se filter + column ek saath
Example
print(df.loc[df["category"] == "Grocery", ["product", "quantity", "price"]])
Output
product quantity price
4 Rice 5kg 2 420.0
6 Tea 1kg 5 380.0
8 Oil 1L 3 160.0
Naya column aur values update karna
Example
df["revenue"] = df["quantity"] * df["price"]
print(df[["product", "quantity", "price", "revenue"]])
Output
product quantity price revenue
0 Headphones 2 1500.0 3000.0
1 T-Shirt 3 499.0 1497.0
2 Jeans 1 1299.0 1299.0
3 Mouse 4 650.0 2600.0
4 Rice 5kg 2 420.0 840.0
5 Keyboard 1 NaN NaN
6 Tea 1kg 5 380.0 1900.0
7 Jacket 1 2499.0 2499.0
8 Oil 1L 3 160.0 480.0
9 Charger 2 899.0 1798.0
Keyboard ka price missing tha, isliye uska revenue bhi NaN aaya. Maan lijiye humne pata kiya ki price 1199 tha — loc se value update karte hain:
Example
df.loc[df["order_id"] == 1006, "price"] = 1199
df["revenue"] = df["quantity"] * df["price"]
print(df.loc[df["order_id"] == 1006, ["product", "price", "revenue"]])
print("Total revenue:", df["revenue"].sum())
Output
product price revenue
5 Keyboard 1199.0 1199.0
Total revenue: 17112.0
Values update karte waqt hamesha
df.loc[condition, column] = valuelikhiye.df[condition][column] = valuejaisa chained assignment kabhi-kabhi original DataFrame ko badalta hi nahi.