Back to blog

The legality of web scraping in India in 2026 and how an MSME can build a defensible lead-generation scraper

Is Web Scraping Legal in India? 2026 Rules for MSME Owners

Cybiqon Team
19 min read
web scrapingDPDP ActlegalMSMEIndialead generation
Is Web Scraping Legal in India? 2026 Rules for MSME Owners

Is Web Scraping Legal in India? 2026 Rules for MSME Owners

Every MSME owner who asks us "is web scraping legal in India?" has been scared by the same number: ₹250 crore. It shows up in the first line of nearly every Indian compliance blog on the subject. So we went and read the actual Schedule to the DPDP Act 2023. That ₹250 crore ceiling attaches to failing to take reasonable security safeguards over data you already hold — under Section 8(5). It does not attach to scraping. A scraping or lawful-basis breach falls into the residual entry at the bottom of the Schedule: up to ₹50 crore. And Section 3 — the section everyone quotes — is not even in force yet.

That gap between the fear and the law is why so many lead-generation projects stall for months at "is this even legal?" while a competitor who never asked ships it. This post fixes that. You will get a straight answer to the question, a decision table showing what you can and cannot build in India today, the six build rules you can hand a developer as a spec, and the two deadlines that can actually hurt you this quarter — neither of which is the one you have been reading about.

The Short Answer: Is Web Scraping Legal in India in 2026?

Web scraping is not illegal in India. There is no Indian statute that prohibits it, no Indian equivalent of the American CFAA prosecutions you will find dominating global search results, and no court in India has held that reading a public web page with a script is unlawful in itself.

What is regulated is what you take, who it belongs to, and what you do next. The legality of data scraping under Indian law turns on four questions, and you can answer all four about your own project in ten minutes:

  1. Is it personal data or entity data? A company name, category, city, GSTIN, landline and a generic info@ address describe a business. A named individual's personal mobile number and personal email describe a person. Only the second category creates real exposure.
  2. Did you agree to a contract to get it? If your scraper never logs in and never clicks "I Accept", the contract claim against you is materially harder to run.
  3. Are you copying facts, or someone's original compilation? Facts are free. A curated, selected, arranged database may be someone's copyright.
  4. What do you do with it afterwards? This is where most Indian MSMEs actually get into trouble — republishing, reselling, or cold-calling scraped numbers.

Get those four right and you are in defensible territory. Get them wrong and the risk is real, but it is almost never the risk you were told about.

Cybiqon builds software, not legal opinions — this is general information as of August 2026, and you should take advice from a qualified Indian lawyer before you scale a scraper that touches personal data.

What You Can Build vs What You Can't: The Decision Table

This is the page most owners actually need. Find the row that matches what you want, and you have your answer.

What you want to build Buildable in India today? The catch
Pull company name, category, city, GSTIN, landline and generic info@ addresses from a public B2B directory into your CRM Yes — generally defensible Entity data, not personal data. Still obey robots.txt and rate-limit.
Collect named individuals' personal mobile numbers and personal email IDs at scale High risk — don't Builds a personal-data corpus you must defend from around May 2027, with no grandfathering for what you already hold.
Scrape Google Maps listings into a stored database No Maps Platform Terms expressly prohibit exporting, extracting or scraping Maps Content and bulk-downloading places information, including copying and saving business names, addresses or user reviews.
Cache Google Maps data for a live feature Only within limits Place IDs are the only field you may store indefinitely. Latitude/longitude may be cached for a maximum of 30 consecutive calendar days, then must be deleted.
Log in behind a click-wrap "I Accept" and then scrape High risk You have formed an electronic contract, and click-wrap terms are enforceable in India where notice is reasonable and acceptance affirmative.
Republish or resell the compiled listings you scraped No This is precisely what got Padawan Ltd permanently restrained in OLX's Delhi High Court suit (order dated 15 December 2016).
Cold-call or SMS scraped numbers from staff personal mobiles No TCCCPR non-compliance leads to a 20-call/20-SMS daily cap, disconnection, and blacklisting for up to two years.
A Chrome extension that scrapes the page a user is on and quietly sends it to your server for enrichment or resale No — and this one is live now Chrome Web Store Limited Use enforcement began 1 August 2026 and explicitly names scraped content.
Scrape competitors' public product and price pages for your own market research Yes — generally Rate-limit, respect robots.txt, and never republish the compilation.
Scrape content to train an AI model Unsettled The Delhi High Court has engaged with this once, on an interim basis only. See below.

If your project sits in a "Yes" row, stop worrying and start building. If it sits in a "No" row, the honest answer is that there is usually a defensible version of the same business outcome one row up — that is the redesign work worth paying for.

The ₹250 Crore Myth: What the DPDP Schedule Actually Says

The Schedule to the Digital Personal Data Protection Act 2023 has seven entries, and they are not interchangeable:

  • Section 8(5) — failure to take reasonable security safeguards: up to ₹250 crore
  • Section 8(6) — failure to notify a personal data breach: up to ₹200 crore
  • Section 9 — obligations regarding children's data: up to ₹200 crore
  • Section 10 — additional duties of a Significant Data Fiduciary: up to ₹150 crore
  • Section 15 — duties of a Data Principal: up to ₹10,000
  • Section 32 — voluntary undertaking
  • Breach of any other provisionup to ₹50 crore

A firm that scrapes personal data without a lawful basis lands in the last entry, not the first. Quoting ₹250 crore at a 20-person firm in Ludhiana for scraping is scaremongering, and it is factually wrong.

Here is the part worth internalising though: ₹250 crore becomes relevant the moment you store scraped personal data carelessly and it leaks. A shared Excel sheet of 14,000 contacts sitting on one sales laptop plus a WhatsApp-forwarded copy, with no deletion policy — that is a Section 8(5) exposure, not a scraping exposure. The number people fear is real; they have simply attached it to the wrong step in their own process. Broader DPDP Act compliance for MSMEs is mostly about how you hold data, not how you got it.

When Does the DPDP Act Actually Bite? Not Yet — and the Board Doesn't Exist

This is the single biggest thing missing from every other page ranking for web scraping laws in India 2026.

Rule 1(2) of the DPDP Rules 2025 (notified 13 November 2025) staggers commencement in three phases. Rules 1, 2 and 17–21 took effect on publication. Rule 4 lands at twelve months, around 13 November 2026. Rules 3, 5–16, 22 and 23 land at eighteen months — on or about 13 May 2027. The companion MeitY notification brings Act Sections 3–5, 6(1)–(8) and (10), 7–17 and 28–34 into force only in that third phase.

Section 3 is the section that carries the DPDP Act publicly available data exemption, the consent architecture and the route to the penalty Schedule. Until it commences, none of it is operative law. (Write the date as "on or about" — professional sources differ by a day, because the Rules were signed on 13 November 2025 and gazetted on 14 November 2025.)

And there is a second gap. The Data Protection Board of India was established in law in November 2025. MeitY issued nomination calls in May and June 2026, with a Search-cum-Selection Committee chaired by the Cabinet Secretary. As reported by LiveLaw on 1 August 2026, the Board still has zero appointed Chairperson and zero Members. There is currently no body in India that can receive a DPDP complaint, hold a hearing or pass a penalty order.

Read that as a build window, not a licence. There is no grandfathering for personal data you are already sitting on when Section 3 switches on. Every month you spend accumulating personal mobile numbers is a month of cleanup work you are buying yourself in 2027. You have roughly nine months to build it correctly the first time — and given that only 2.5% of MSMEs say they understand the DPDP Act (India SME Forum's "Readiness to Comply" study across around one lakh members, reported September 2025), most of your competitors will not use them. PwC India's 2024 survey found similar: 16% of Indian consumers and just 9% of Indian organisations claimed comprehensive understanding of the Act.

The Law That Actually Applies Today: IT Act, Copyright and Contracts

Since Section 3 is not live, here is the live law in August 2026.

IT Act 2000, Section 43 creates civil liability where a person acts "without permission of the owner or any other person who is incharge of a computer, computer system or computer network" and, among other things, "downloads, copies or extracts any data, computer data base or information". Note the hinge word: permission. The section as it now stands states no monetary ceiling — the widely repeated "₹1 crore cap" is the pre-2008 position and is no longer the law. This is the core of IT Act Section 43 web scraping exposure in India.

IT Act 2000, Section 66 converts a Section 43 act into a criminal offence only where it is done dishonestly or fraudulently. That mens rea requirement is doing a lot of work. Penalty: up to three years' imprisonment and/or a fine up to ₹5 lakh. For context, the big Indian data cases that end in arrests involve theft and resale of breached databases — not scripts reading public pages. Conflating the two is exactly the error this post is pushing back on.

Copyright Act 1957, Section 2(o) includes "tables and compilations including computer databases" within "literary work". India has no sui generis database right, so Copyright Act 1957 database protection in India runs entirely through this route. Two cases set the test:

  • Burlington Home Shopping v. Rajnish Chibber (Delhi HC, 1995) — a compiled customer database was protectable as a literary work, with originality measured by skill, labour and judgment in selection and arrangement.
  • Eastern Book Company v. D.B. Modak, (2008) 1 SCC 1 — the Supreme Court's "modicum of creativity" standard. Raw public-domain facts are not protected; original selection, arrangement and editorial additions are.

So: a phone number is a fact. Someone's curated, categorised, verified directory of 8,000 of them may be their literary work. Copy the facts you need; don't clone the compilation.

Click-wrap vs browse-wrap. Click-wrap terms — where notice is reasonable, terms are accessible before acceptance, and there is an affirmative "I Accept" — are enforceable; IT Act Sections 4 and 10A recognise electronic contracts. Browse-wrap, a terms hyperlink buried in a footer, is materially weaker and needs actual or constructive notice. The practical upshot for is-scraping-IndiaMART-data-legal type questions: a scraper that never creates an account and never clicks "I Agree" is in a very different position from one that does.

Does robots.txt have legal force in India? No — robots.txt is not itself a contract and there is no Indian decision making it binding. But ignoring it is the cleanest available evidence of bad faith, and a court weighing "without permission" under Section 43 will look at it. Obey it anyway; it is free.

MeitY's stated position. As reported from Minister of State Jitin Prasada's Rajya Sabha reply of 18 February 2025 (and an earlier answer in August 2024), "web scraping of any publicly available user data by any intermediary... is regulated under the Information Technology Act, 2000", with reference to Section 43 and the IT Rules 2021. Read precisely, that is not a denial that the DPDP publicly available data exemption exists — it is a statement that scraping is separately regulated by the IT Act regardless of the DPDP carve-out. Two statutes doing two different jobs. The signal for you: the exemption is not a general licence.

What about AI training? In late July 2026 the Delhi High Court refused an interim injunction in ANI Media Pvt. Ltd. v. Open AI OpCo LLC, neutral citation 2026:DHC:5900 (Justice Amit Bansal) — India's first substantive judicial engagement with AI training and scraping. On a prima facie basis the court indicated that storage for training could fall within Section 52(1)(a)(i) of the Copyright Act, that Parliament deliberately omitted "non-commercial" from that sub-clause, that ChatGPT is not a market substitute for a news syndication business, and that Indian courts do have territorial jurisdiction even where training and storage happened on US servers, because output is reproduced in India. This is an interim order on a prima facie basis, expressly not binding on final adjudication, and the suit is still pending trial. Nobody legalised scraping in July. Separately, MeitY's India AI Governance Guidelines of November 2025 recommended that the DPIIT committee revisit Section 52 to consider a text-and-data-mining exception — the government itself acknowledging the law here is unclear.

Six Build Rules You Can Hand Your Developer

Picture a 22-person auto-components manufacturer in Ludhiana. Two junior sales executives spend most of the week copy-pasting supplier and buyer listings off IndiaMART, Justdial and Google Maps into a shared Excel sheet, then cold-calling every number from personal mobiles. The sheet holds names, personal mobile numbers and email IDs at around 14,000 firms. It sits on one laptop plus a WhatsApp copy, has no deletion policy, and nobody has asked whether any of it is lawful. That is the composite we see constantly.

Here is the defensible rebuild — six rules, every one verifiable today. Paste them into your developer brief:

  1. Scrape firmographics, not people. Company name, category, city, product listings, GSTIN, landline, generic info@ address. Entity data does not identify an individual, which is the test that matters under DPDP Section 2(t).
  2. Drop named-individual mobile numbers and personal emails at ingestion, not in a cleanup later. If the field never enters your database, there is no personal-data corpus to defend.
  3. Do not scrape Google Maps for a stored database. The Maps Platform Terms are explicit. Place IDs are the only field storable indefinitely; latitude/longitude may be cached for a maximum of 30 consecutive calendar days. If you want map coverage of your own business instead, that is a local SEO problem, not a scraping one — and it is exactly why your competitor shows up on Google and you don't.
  4. Read and obey robots.txt, and rate-limit hard. Not because robots.txt binds you, but because ignoring it is the evidence a court will use against you. Reference your own crawl policy in your terms of service too.
  5. Never re-publish or resell the compiled listings. Use them internally. The moment you put someone else's compilation back on the internet, you are in OLX v. Padawan territory.
  6. Route all outbound calling and SMS through a registered telemarketer path on DLT, with 140/1600-series headers.

The honest outcome: your sales team stops doing data entry and starts making calls, your outbound list becomes an asset you can show a customer or an acquirer, and no quiet personal-data liability accumulates on a laptop before May 2027. This is also the point at which scraping becomes genuine market research rather than a bigger call list — pricing movements, competitor catalogue changes, distributor coverage gaps.

The Two Deadlines Nobody Is Watching

While everyone stares at a 2027 penalty regime with no Board to enforce it, two things can hit you this quarter.

1. Chrome Web Store Limited Use — enforcement began 1 August 2026. Google's Limited Use policy explicitly extends to "scraped content or otherwise automatically gathered user data" and requires that all collected data be strictly necessary to the extension's single disclosed purpose. It flatly prohibits transferring or selling user data to third parties such as data brokers. Non-compliant extensions can be removed from the store. That directly targets the commonest MSME extension pattern: scrape the directory page the user is already on, and also quietly ship it to your server for enrichment or resale. The Chrome extension data scraping policy for 2026 is not a future risk — it started 17 days before this post went up. If you are building a tool for pulling verified Indian business contacts, scope it to one disclosed purpose and keep the data on the user's machine.

2. TCCCPR. This is the regulation most likely to actually punish an Indian MSME that scrapes numbers and cold-calls them. Non-compliance leads to a usage cap of 20 calls and 20 SMS per day, then disconnection, then blacklisting for up to two years. TRAI has already blacklisted over 1,150 entities and disconnected 18.8 lakh telecom resources. Losing your sales team's numbers for two years is a far more concrete threat than a penalty Schedule that is not in force.

One more piece of context worth having: automated traffic accounted for 53% of all web traffic in 2025, up from 51% in 2024, according to Imperva's 2026 Bad Bot Report. More than half the requests hitting your own website are already bots. Your competitors are scraping you, and sites are blocking harder — which is why polite, rate-limited, identifiable scrapers survive and aggressive ones get banned within a week.

FAQs

Is web scraping illegal in India?

No. There is no Indian law that makes web scraping illegal in itself. Liability arises from context: taking data "without permission" under IT Act Section 43, copying someone's original compilation under the Copyright Act 1957, breaching a click-wrap contract you accepted, or misusing personal data. Scraping public entity data politely, for internal use, is generally defensible.

Does the DPDP Act 2023 apply to publicly available data?

Section 3(c)(ii) says the Act does not apply to personal data "made or caused to be made publicly available by (A) the Data Principal to whom such personal data relates; or (B) any other person who is under an obligation under any law for the time being in force in India to make such personal data publicly available." Note the limits: the individual must have published it themselves, or a law must have required its publication. A number scraped from a directory the person never chose to publish on is not obviously covered — and MeitY's position is that scraping is separately regulated under the IT Act 2000 anyway.

Can I be fined ₹250 crore for scraping data in India?

Not for scraping. The ₹250 crore ceiling in the Schedule to the DPDP Act attaches to a Section 8(5) failure to take reasonable security safeguards. A scraping or lawful-basis breach falls under the residual entry, capped at ₹50 crore. Both are academic until Section 3 commences on or about 13 May 2027, and as of 1 August 2026 the Data Protection Board had no Chairperson and no Members appointed.

Is it legal to scrape leads from IndiaMART, Justdial or Google Maps?

Different answers for different sites. Directory listings of firmographic data, scraped without logging in and without accepting click-wrap terms, are the most defensible case. Google Maps is the clearest "no": the Maps Platform Terms expressly prohibit extracting, scraping or bulk-downloading Maps Content, including saving business names, addresses and reviews. Whether scraping IndiaMART data is legal depends heavily on whether you created an account and clicked "I Accept" first.

Is robots.txt legally binding in India?

No Indian court has held robots.txt to be contractually binding, and it does not have independent legal force here. But it is strong evidence of the site owner's stated permission — and "without permission of the owner" is the exact test in IT Act Section 43. Ignoring robots.txt is free evidence handed to the other side. Obey it.

Do I need consent to cold-call numbers I scraped from a public directory?

The bigger constraint today is not consent law but telecom regulation. Commercial calling and SMS in India must run through a registered telemarketer path on DLT with 140/1600-series headers. Non-compliance means a 20-call/20-SMS daily cap, disconnection, and blacklisting for up to two years. From May 2027, DPDP obligations sit on top of that.

Work With Cybiqon on the Build, Not Just the Question

Cybiqon AI Solutions builds websites, apps, Chrome extensions and AI automation for Indian MSMEs — and because we build the scraper, the browser extension and the AI enrichment layer as one owned system, we can architect the legality into the build itself: firmographic-only fields, no personal-data storage, robots.txt and rate-limit respect, ToS-safe sources, Chrome Limited Use single-purpose scoping, and DLT-compliant outbound.

That also means we will tell you honestly when something isn't buildable, and propose the version that is. Most owners find the compliant design gets them 90% of the commercial outcome with none of the 2027 cleanup.

Tell us what you want to scrape and we will give you a straight answer. Visit cybiqon.in, call +91 9250711473, or write to [email protected].

The Takeaway

So, is web scraping legal in India? Yes, in the shape most MSMEs actually need it — public firmographic data, taken politely, kept internal, never resold. The ₹250 crore number you were scared of belongs to a different section and a regime that switches on around 13 May 2027, in front of a Board that does not yet exist. Meanwhile the Chrome Web Store and TCCCPR can reach you this month. Build to the six rules now, while the window is open — and if you want a second pair of eyes on the design, Cybiqon is a call away.

Want this set up for your business?

Book a free call — no tech jargon, no sales pressure. Just honest answers.

WhatsApp us
Chat with us!