WSWhat Scene?

Research · 8 min read

The invisible business problem

We surveyed 9,779 Pune businesses recorded in OpenStreetMap and checked every website on record. 82.1% have no independent web presence at all, and 10.1% own a website. Of the whole survey, 4.9% had a site that was working and current.

What we did

We pulled every business in OpenStreetMap across the Pune area that matched our category selectors, 9,779 of them, classified each into a sector, and then fetched and inspected every website that was on record. The result is a picture of how much of a real Indian city's commerce is reachable online, measured rather than estimated.

The count is a floor rather than a census, and the selectors are part of why. A mapped business whose tags fall outside the categories we asked for is not in the total at all, so this understates rather than overstates. The limits section has the rest.

We are publishing the statistics and not the dataset. That is partly a licence question and partly a straight commercial one, and both are covered further down. Every figure in the tables below is read from a recorded run rather than typed into the page.

The headline

82.1% of mapped Pune businesses (8,033 of 9,779) have no independent web presence at all: no website, no social page, and no phone number recorded against them either. 10.1% own a website.

PresenceCountShare
Offline only (invisible)8,03382.1%
Own website98910.1%
Listing/phone only (no site)7317.5%
Social only (no site)260.3%
Digital presence across the survey

The gap between those two figures is where the interesting part sits. A business with a listing on an aggregator is findable, but on someone else's terms, in someone else's ranking, next to its competitors, with the relationship owned by the platform. That is a different position from having nothing, and a different position again from owning the surface people find you on.

Having a website is not the same as having a working one

We fetched every site on record rather than counting the ones that existed. That distinction turns out to matter.

StatusCountShare
NO_WEBSITE8,76489.6%
MODERN_OK4834.9%
OUTDATED2432.5%
BROKEN2052.1%
PARKED350.4%
SOCIAL_ONLY320.3%
SHOPIFY170.2%
Website status as recorded by the audit

Across the entire survey, 483 businesses (4.9%) had a site the audit classified as working and current. Every other record fell into one of the remaining buckets in the table: NO_WEBSITE, OUTDATED, BROKEN, PARKED, SOCIAL_ONLY, SHOPIFY.

The two tables classify slightly differently, so their categories do not sum to each other. Presence is about what a business has; status is about what its site did when we fetched it. We are publishing both as recorded rather than reconciling them into one tidier number that neither column actually supports.

The denominator that changes the finding

The table above is a share of all 9,779 businesses, which is the right denominator for asking how many are online and the wrong one for asking what condition these sites are in. Against the whole survey a broken site looks like a rounding error, because it is being averaged against thousands of businesses that have no site to break.

Measured instead against the 983 businesses that actually have a website, the same records read very differently. Social-only records are excluded from this base: a social page is not a website in any condition, and counting it would pad the denominator and understate the failure rate.

ConditionCountShare of sites
MODERN_OK48349.1%
OUTDATED24324.7%
BROKEN20520.9%
PARKED353.6%
SHOPIFY171.7%
Condition among the 983 businesses that have a website

20.9% of the websites in this survey were broken when we fetched them, and 49.1% were working and current. The same records, against the population rather than against the sites, put broken at 2.1%. Neither figure is wrong. They answer different questions, and only one of them is a statement about websites.

This is the most quotable pair of numbers on the page and the easiest to misuse, so both are stated together on purpose. Anybody citing the failure rate should say which base it is against.

What the sites that exist are built on

Platform was detected on 273 of the 983 sites, which is 27.8% coverage. Every percentage in the table below is a share of those 273, not of all sites and not of all businesses.

PlatformCountShare of detected
wordpress22783.2%
wix176.2%
shopify176.2%
squarespace51.8%
godaddy-builder51.8%
weebly20.7%
Platform, among the 273 sites where one was detected

wordpress accounts for 83.2% of the sites where a platform was identified, which is worth reading with the coverage figure attached rather than as a statement about the whole city.

Undetected does not mean hand-built. It means the probe could not tell, which is a different statement, and it applies to 72.2% of the sites. Treat this table as a floor on the share each platform holds rather than as a measurement of the market.

By sector

The two largest sectors account for 59.2% of the survey between them, which says as much about what OpenStreetMap records well, and about which categories we asked for, as it does about Pune.

SectorCountShare
Food & Hospitality2,92329.9%
Medical & Healthcare2,87029.3%
Retail & Shopping1,16811.9%
Education9189.4%
Other Local Services4734.8%
Sports & Fitness3403.5%
Automotive3283.4%
Beauty & Wellness2392.4%
IT & Technology1331.4%
Home & Construction1311.3%
Events & Entertainment800.8%
Professional Services760.8%
Nonprofits & Trusts420.4%
Travel & Transport380.4%
Manufacturing & Industrial200.2%
Businesses by sector

We are not publishing the sub-sector breakdown. It is the most useful cut in the dataset and it is also a targeting map, and the finding on this page does not need it.

The number we deliberately do not lead with

The single most quotable figure in this dataset is that 89.6% of records are marked NO_WEBSITE. It is arithmetically correct and we will not use it as a headline, because of what it actually means.

NO_WEBSITE means there is no website tag in OpenStreetMap. That is an absence of data, not a verified absence of a website. Those are different claims, and treating the first as the second is the most common way research goes wrong: a gap in a source becomes a fact about the world somewhere between the spreadsheet and the headline.

So the headline is built on the presence classification, and every figure here is qualified as being about businesses as recorded in OpenStreetMap. The weaker claim is the one the data supports.

If you take one thing from this page and you are not in Pune, take that. A dataset's silence is not evidence. It is the thing you have to go and check, and the check is usually the expensive part everybody skips.

Method

  • Source: OpenStreetMap, covering the Pune district. Contributors licence it under ODbL.
  • Every record was classified into a sector and sub-sector from its category tag and its business name.
  • Every website on record was then fetched and inspected, and classified as working and current, outdated, broken, parked, or a social page in place of a site.
  • Records with no website tag were counted as such and not probed further, because there was nothing to probe.
  • Aggregates are recomputed from the source by a script, and this page reads its output directly rather than restating it.

We are describing the method rather than publishing the queries, the classification keyword lists or the tooling. That is enough for a reader to judge how much weight the numbers carry, which is what a method section is for. It is not enough to rebuild the dataset, which is deliberate.

What this does not show

The limits are in their own section below and they are worth reading before quoting anything here. The short version: this measures what one map records about one city at one point in time, and every figure inherits that.

Why we publish the statistics and not the data

OpenStreetMap is licensed under the Open Database Licence. Publishing statistics derived from it requires attribution. Publishing a derived database would require us to license that database share-alike, which would mean handing the whole thing to anyone who wanted it, including anyone competing with us.

There is a plainer reason as well, and it would apply without the licence. Building this took real work and the result is a commercial asset. Publishing the finding costs us nothing and publishing the file would cost us the asset.

An earlier draft of our own publication rule allowed a per-row export that kept only sector, web status, coordinates and the OpenStreetMap node id, reasoning that none of those is a contact detail. That was wrong. A node id resolves directly to a business name and address on openstreetmap.org, so the file would have been the list with one join step in front of it. No phone column is not the same thing as not reconstructable, and coordinates and source ids are identifiers whatever else they are.

What this does not cover

  • These figures describe businesses as recorded in OpenStreetMap, not all businesses in Pune. Coverage varies by area and by sector, and a business that is not mapped is not in the survey at all.
  • The survey covers what our category selectors asked for. A mapped business whose tags fall outside those categories is missing from the total, so the count is a floor set partly by the query rather than a census.
  • The area was taken as a rectangle around Pune rather than the administrative boundary, so its corners reach into neighbouring districts. Those corners are rural and the effect is small, but it is not zero.
  • The offline-only category means no website, no social page and no phone number recorded in OpenStreetMap. It does not mean we checked a directory or a listing site and found nothing. We did not.
  • NO_WEBSITE means no website tag in OpenStreetMap. It is an absence of data, not a verified absence of a website, which is why it is not the headline.
  • Large institutions such as colleges and major hospitals are the least reliable category. OpenStreetMap frequently lacks a site they do in fact have, so their absence is overstated here.
  • A site returning HTTP 403 or 429 was recorded as broken, and those codes often mean the site is blocking automated requests rather than being down. Broken is therefore an overcount.
  • The survey includes national chains with a Pune outlet, so not every record is an independent local business.
  • This is one city at one point in time. Nothing here establishes a trend, and nothing here transfers to another city without measuring it.
  • We have not contacted any of these businesses to verify their status, and a survey by fetching is not the same as a survey by asking.

Sources

  • Pune business survey, 9,779 records sourced from OpenStreetMap across the Pune district. Sector classification from category tags and business names, followed by a fetch and inspection of every website on record. Aggregates recomputed from the source at publication time. Run: scripts/pune-aggregates.mjs, 2026-08-30 14:14 UTC.
  • OpenStreetMap copyright and licence, OpenStreetMap Foundation. Retrieved 23 August 2026.

Revisions

  • 30 August 2026 Added the condition of the sites that exist and the platform breakdown. Same dataset, recomputed against the businesses that actually have a website rather than against the whole survey.
  • 23 August 2026 First published, with all aggregates recomputed from the source dataset by script.

This page is revised in place rather than replaced, so its address does not change.

Next step

Want this built, not just explained?