
You already know lists (Parts 17-18) and tuples (Part 19). All three are collections, but they answer different questions.
| List | Tuple | Set | |
|---|---|---|---|
| Main question it answers | What is at position i? | What is at position i in a fixed row? | Does this value already exist? |
| Order / index | Yes ([0], slicing) | Yes | No - no indexing |
| Duplicates | Allowed | Allowed | Not stored - only unique members |
| Mutability | Mutable | Immutable | Mutable |
| Best mental model | Working sequence / basket | Locked row / fixed record | Membership container |
A set is not "a list without duplicates." A set is a different kind of container made for a different job.
That is why converting a list to a set is common:
user_ids = [101, 102, 103, 101, 104, 102]
unique_ids = set(user_ids)
print(unique_ids) # {101, 102, 103, 104}
A set is an unordered collection of unique values.
(Technically, only hashable values can be stored — we explain what "hashable" means in the "How Set Works Internally" section below.)
skills = {"Python", "SQL", "Docker"}
| Property | Meaning |
|---|---|
| Unordered | No position contract, so no indexing or slicing |
| Unique | Equal values collapse into one member |
| Mutable | You can add and remove members |
Unordered does not mean random.
It means Python does not promise a usable position like 0, 1, 2.
That is why this has no meaning:
skills[0] # TypeError
A set is a membership container, not a numbered sequence.
A set exists because sometimes we do not care about order. We care about two different questions:
A list can answer those questions too, but it answers them by scanning one by one. A set answers them using hashing, which is much faster on average for lookup.
# Conceptually:
# List membership: item in big_list → scans one by one (slow for large data)
# Set membership: item in big_set → hash-jump to the right area (fast on average)
So the real story is:
This is your first strong bridge to DSA thinking.
# Curly braces with members
colors = {"red", "green", "blue"}
# Duplicate removal happens automatically
numbers = {1, 2, 2, 3, 3, 3}
print(numbers) # {1, 2, 3}
# From a list
names = ["Asha", "Ravi", "Asha", "Priya", "Ravi"]
unique_names = set(names)
print(unique_names)
# From a string
letters = set("banana")
print(letters) # {'b', 'a', 'n'} in some order
wrong = {}
print(type(wrong)) # <class 'dict'>
correct = set()
print(type(correct)) # <class 'set'>
{} creates an empty dictionary, not a set.
Always use set() for an empty set.
Because a set has no position contract, indexing is not supported.
colors = {"red", "green", "blue"}
print(colors[0])
# TypeError: 'set' object is not subscriptable
With a list, colors[0] means:
With a set, there is no user-facing slot 0. The internal placement is based on hash math, not user-visible position.
x[0], slicing, or stable order? Use list or tuple..add(x) - add one memberskills = {"Python", "SQL"}
skills.add("Docker")
print(skills)
If the member already exists, nothing breaks and nothing new is added:
skills.add("Python")
print(skills) # still same set
.remove(x) - strict removeskills = {"Python", "SQL", "Docker"}
skills.remove("SQL")
print(skills)
skills.remove("Java")
# KeyError
.discard(x) - safe removeskills = {"Python", "Docker"}
skills.discard("Java") # no error
print(skills)
.pop() - removes an arbitrary memberskills = {"Python", "SQL", "Docker"}
removed = skills.pop()
print(removed)
print(skills)
Do not explain .pop() like list pop-from-end.
For a set, .pop() removes an arbitrary member.
.clear() - remove everythingskills = {"Python", "SQL"}
skills.clear()
print(skills) # set()
This is where sets become visually powerful and mathematically clean.
frontend = {"HTML", "CSS", "JavaScript", "React"}
backend = {"Python", "JavaScript", "SQL", "Docker"}
| - all unique members from bothall_skills = frontend | backend
print(all_skills)
all_skills = frontend.union(backend)
& - common memberscommon = frontend & backend
print(common) # {'JavaScript'}
common = frontend.intersection(backend)
- - first set minus secondonly_frontend = frontend - backend
print(only_frontend)
only_backend = backend - frontend
print(only_backend)
^ - in either, but not bothexclusive = frontend ^ backend
print(exclusive)
frontend: {HTML, CSS, JavaScript, React}
backend: {Python, JavaScript, SQL, Docker}
Union (|): everything unique from both
Intersection (&): only shared members
Difference (-): direction matters
Symmetric (^): everything except the shared part
These operations usually build a new set object.
So this is a good place to connect with id():
a = {1, 2, 3}
b = {3, 4, 5}
c = a | b
print(id(a))
print(id(c)) # different id - new object
That matches earlier parts:
The real superpower of a set is fast average-case membership testing.
big_list = list(range(1_000_000))
big_set = set(range(1_000_000))
print(999_999 in big_list)
print(999_999 in big_set)
Both return True, but they work very differently.
For a list, Python checks one member after another:
That is a scan. Average idea: O(n).
For a set, Python:
Average idea: O(1).
Hash narrows the search; equality confirms the match.
in Is FastThis is the crux section. Remember Part 6 — your Python code becomes bytecode, and the PVM executes it in C. The hash table inside a set is part of that C layer. That is where the speed comes from.
A set can only store hashable objects.
Noneboolintfloatcomplexstrbytesrangetuple - only if all items inside are hashablefrozenset - only if all items inside are hashablelistdictsetbytearrayMutable built-in containers are not hashable. Most immutable built-ins are hashable.
print(hash(42))
print(hash("hi"))
print(hash((1, 2)))
print(hash([]))
# TypeError: unhashable type: 'list'
That is exactly why this fails:
bad = {[1, 2]}
# TypeError: unhashable type: 'list'
And this works:
good = {(1, 2)}
print(good)
Do not teach this as only: "immutable = hashable"
The real rule is:
For now, the practical shortcut covers 99% of cases: if it is a mutable built-in container (list, dict, set), it is not hashable. If it is immutable (int, str, tuple, frozenset), it usually is.
hash(x) vs id(x)Students must not confuse these two.
| Function | Meaning |
|---|---|
id(x) | identity of the object |
hash(x) | fingerprint integer used for hash-based lookup |
x = "python"
print(id(x))
print(hash(x))
These numbers are different because they do different jobs.
id(x) = who this object ishash(x) = where to start looking in a hash-based containerFor strings and bytes, hash values can differ across process runs because Python enables hash randomization by default. Inside one run they stay consistent, but across runs they may change.
in on list vs in on settarget in my_list
Conceptually:
This is why large lists become slower for repeated membership checks.
target in my_set
Conceptually:
hash(target)This is why membership is fast on average.
A list finds by walking. A set finds by jumping.
Inside the set object, CPython maintains an internal table of slots.
Important: this is not a normal Python list that you can access. It is a lower-level implementation detail hidden inside the set object.
name on stack/frame ---> set object in heap
|
|-- hidden table of slots
|-- each occupied slot stores:
- cached hash code
- reference to member object
You, as a Python programmer, do not control these internal slot numbers. They are for the runtime, not for user-facing indexing.
That is why:
s[0]When a member is inserted:
hash(member)You do not need to reimplement this in this part. Students only need the idea:
hash -> slot selection -> possible probe -> equality confirmation
A collision means two different members want the same slot area.
Python handles this internally. It does not give up, and it does not scan the whole structure like a list. It follows a probe sequence to check nearby slots.
Say:
Do not go deep into:
That belongs in a separate DSA episode.
A hash table cannot stay packed tight forever. It needs empty space to stay fast.
So conceptually:
Use sys.getsizeof() just as evidence that the set grows in jumps:
import sys
s = set()
for i in range(20):
s.add(i)
print(i, sys.getsizeof(s))
sys.getsizeof() shows the size of the set object itself, not the full recursive size of every object it references.
CPython starts with a small internal table (commonly 8 slots) and grows as needed. The exact starting size and growth strategy are implementation details — the CPython team keeps them internal so they can optimize performance freely between Python versions without breaking anyone's code.
hash is not password hashingThis is a one-sentence cleanup point.
The word hash is used in two very different worlds:
hashlib, bcrypt, SHA, etc.)Same English word, different purpose.
This is the moment to make students feel powerful.
DSA asks two questions:
That means students have already started learning DSA thinking.
id()skills = {"Python", "SQL"}
print(id(skills))
skills.add("Docker")
print(id(skills)) # same id
The object identity stays the same. That means the same set object was changed in place.
a = {1, 2}
b = {2, 3}
c = a | b
print(id(a))
print(id(c)) # different id
So:
.add(), .remove(), .discard(), .clear(), .update() = mutate same seta | b, a & b, a - b, a ^ b = build new set objects| Need | Best choice | Why |
|---|---|---|
| Order matters | List / Tuple | position is meaningful |
| Need indexing | List / Tuple | set has no indexing |
| Need duplicates | List / Tuple | set removes duplicates |
| Need fixed row | Tuple | immutable sequence |
| Need uniqueness | Set | duplicates collapse |
| Need fast average membership | Set | hash-based lookup |
A list answers: "What is at index i?"
A set answers: "Does x exist?"
| Method | Meaning |
|---|---|
.add(x) | Add one member |
.remove(x) | Remove member, raise KeyError if missing |
.discard(x) | Remove member, no error if missing |
.pop() | Remove and return an arbitrary member |
.clear() | Remove all members |
.copy() | Shallow copy |
| Method / operator | Meaning |
|---|---|
.union(other) or ` | ` |
.intersection(other) or & | Common members |
.difference(other) or - | In first, not in second |
.symmetric_difference(other) or ^ | In either, but not both |
| Method | Meaning |
|---|---|
.update(other) | In-place union |
.intersection_update(other) | Keep only common members |
.difference_update(other) | Remove members found in other |
.symmetric_difference_update(other) | Keep non-common members |
| Method | Meaning |
|---|---|
.issubset(other) | Is every member in other? |
.issuperset(other) | Does this contain all of other? |
.isdisjoint(other) | Do they share nothing? |
frozenset - Immutable Setpermissions = frozenset({"read", "write"})
print(permissions)
permissions.add("delete")
# AttributeError
tuple is to listfrozenset is to setfrozenset exist?Because sometimes you want set behavior, but the container itself must be immutable.
role_permissions = {
frozenset({"read", "write"}): "editor",
frozenset({"read"}): "viewer"
}
dict, set, frozensetThis is your bridge to Part 21.
dict and set are hash-table siblings.
list is a different family.
This prepares students for the next part naturally.
Do not use a set when:
bad = {[1, 2], [3, 4]}
# TypeError: unhashable type: 'list'
good = {(1, 2), (3, 4)}
print(good)
If your app needs to keep insertion sequence for display, do not use a set as the main display container. Use a list or another structure for presentation.
| Operation | Set | Why |
|---|---|---|
x in s | O(1) average | hash-based lookup |
s.add(x) | O(1) average | place by hash |
s.remove(x) | O(1) average | find by hash |
len(s) | O(1) | tracked internally |
| `s1 | s2` | O(len(s1) + len(s2)) |
s1 & s2 | O(min(len(s1), len(s2))) average idea | compare against smaller side |
Say average-case O(1), not absolute magical O(1). Collisions exist, but Python's implementation is designed so average membership stays fast.
This section answers the question students always ask: "Where do we use sets in real work?"
APIs receive data from users, forms, mobile apps. Duplicate entries are common. Before inserting into a database, backends routinely deduplicate.
submitted_emails = [
"ravi@gmail.com", "asha@outlook.com", "ravi@gmail.com",
"dev@yahoo.com", "asha@outlook.com"
]
unique_emails = set(submitted_emails)
print(f"Received {len(submitted_emails)}, unique: {len(unique_emails)}")
This pattern is used in Django, FastAPI, Flask backends every day — whenever form data, webhook payloads, or batch uploads arrive with potential repeats.
Payment gateways like Razorpay or Stripe can send the same webhook event multiple times. Backends track processed event IDs in a set (or Redis set) to avoid double-processing.
processed_events = set()
incoming_events = ["evt_1001", "evt_1002", "evt_1001", "evt_1003", "evt_1002"]
for event_id in incoming_events:
if event_id in processed_events:
print(f"Skipping duplicate: {event_id}")
continue
processed_events.add(event_id)
print(f"Processing: {event_id}")
In production, the processed_events set may live in Redis (which has a native SET data type) rather than Python memory, but the concept is identical.
Web applications check user permissions before allowing actions. A set of permissions makes in checks instant.
user_permissions = {"read", "write", "deploy"}
if "deploy" in user_permissions:
print("Deploy access granted")
else:
print("Access denied")
Frameworks like Django (user.has_perm) and FastAPI (dependency-injected auth) use this pattern internally — checking membership in a collection of granted permissions.
Product teams ask: "Which users are in both the free tier and the waitlist?" or "Which premium users have not completed onboarding?"
free_users = {"Asha", "Ravi", "Nisha", "Dev", "Meera"}
waitlist = {"Dev", "Nisha", "Kiran", "Priya"}
in_both = free_users & waitlist
print(f"In both groups: {in_both}")
only_free = free_users - waitlist
print(f"Free but not on waitlist: {only_free}")
This is intersection and difference — the same math from the set operations section, applied to real user segmentation.
A web crawler or an AI research agent must not revisit the same URL forever. A set of visited URLs prevents infinite loops.
visited_urls = set()
urls_to_crawl = [
"https://example.com/page1",
"https://example.com/page2",
"https://example.com/page1",
"https://example.com/page3",
]
for url in urls_to_crawl:
if url in visited_urls:
print(f"Already visited: {url}")
continue
visited_urls.add(url)
print(f"Crawling: {url}")
This is the same "seen before?" pattern as webhook dedup — applied to crawlers, scrapers, and AI agent loops that explore links or tool calls.
In Retrieval-Augmented Generation (RAG), multiple search queries can return overlapping document chunks. Before stuffing context into the LLM prompt, you deduplicate by document ID.
query_1_results = {"doc_101", "doc_205", "doc_312"}
query_2_results = {"doc_205", "doc_312", "doc_489"}
unique_docs = query_1_results | query_2_results
print(f"Total unique documents for context: {len(unique_docs)}")
Without dedup, the same paragraph would appear twice in the prompt — wasting tokens and confusing the model.
Building a vocabulary from text is a classic NLP preprocessing step. Sets handle uniqueness naturally.
text = "python is great and python is simple"
words = text.split()
vocabulary = set(words)
print(f"Total words: {len(words)}, Unique words: {len(vocabulary)}")
print(vocabulary)
Libraries like NLTK, spaCy, and scikit-learn's CountVectorizer use similar ideas internally — a set (or set-like structure) of unique tokens.
tap_log = ["Majestic", "Indiranagar", "Majestic", "MG Road", "Indiranagar"]
unique_stations = set(tap_log)
print(f"Tapped {len(tap_log)} times, visited {len(unique_stations)} unique stations")
Any system that logs repeated events (metro taps, toll booths, bus stops) and needs distinct counts uses this pattern.
A government office or bank receives applications — the same person may apply multiple times.
applications = [
"XXXX-1234", "XXXX-5678", "XXXX-1234",
"XXXX-9012", "XXXX-5678", "XXXX-3456"
]
unique_applicants = set(applications)
print(f"Total applications: {len(applications)}")
print(f"Unique applicants: {len(unique_applicants)}")
Your shopping basket is still a list because order and history of actions may matter. But if the store manager asks: "How many distinct products touched the scanner today?" — that becomes a set question.
scans = ["rice_5kg", "dal_1kg", "rice_5kg", "oil_1l", "rice_5kg"]
distinct = set(scans)
print(f"{len(scans)} scans, {len(distinct)} distinct SKUs")
In algorithms (BFS, DFS) and AI agent planners, a visited set prevents revisiting the same node.
visited = set()
Whenever your code needs to track "already visited / already processed / already seen", that is a strong set signal. This applies to:
Use this internally while teaching.
A list is like a row of numbered desks in a classroom. You can say:
Position is the meaning.
A set is like the security register at a tech park gate (Manyata, EcoWorld, Electronic City — any tech park your audience knows).
The security guard does not care what order people arrived. The guard only checks:
You do not ask: "Give me the person at position 2." You ask: "Is this person in the register?"
That is why set is not about position. It is about fast yes/no membership logic.
You have enrollment data for two classes:
class_a = ["Asha", "Ravi", "Priya", "Dev", "Meera", "Ravi"]
class_b = ["Ravi", "Dev", "Kiran", "Nisha", "Asha", "Kiran"]
Both classes: {'Asha', 'Ravi', 'Dev'}
Only Class A: {'Priya', 'Meera'}
Only Class B: {'Kiran', 'Nisha'}
All students: {'Asha', 'Ravi', 'Priya', 'Dev', 'Meera', 'Kiran', 'Nisha'}
Exactly one class: {'Priya', 'Meera', 'Kiran', 'Nisha'}
Printed order may differ because sets do not guarantee element position.
Save as src/student_sets.py.
Today you saw a container that answers:
Does this member exist?
Next, we move to a container that answers:
If this key exists, what value belongs to it?
That next container is the dictionary.
Both belong to the same hash-table family.
Comprehensions: FOR LOOP ಬೇಡ, ONE LINE ಸಾಕು! | Python in Kannada | Part-23
Part 23
Comprehensions: FOR LOOP ಬೇಡ, ONE LINE ಸಾಕು! | Python in Kannada | Part-23
Part 23