LESSON 12.2 — Regional Surveys, Cluster/Factor Analysis, I-O Techniques & Spatial Analysis
A. Standard Map
| Topic | Governing Source | Exam Focus |
|---|---|---|
| Regional surveys | Multiple types for plan preparation | Survey types + purpose |
| Patrick Geddes | “Valley Section” + survey–analysis–plan | Foundational logic |
| Cluster analysis | Statistical grouping of similar observations | Method + use |
| Factor analysis | Reduce many variables to fewer dimensions | Method + use |
| Input-Output (I-O) techniques | Sector × sector economic transactions matrix | Method + use |
| Spatial analysis | Density patterns, distribution, land use | Concepts + measures |
| Data collection methods | Primary (surveys, RS, GIS) vs secondary (Census) | Sources |
| Sampling | Size and method | Sample arithmetic |
B. Why It’s Used
Paper II §12 of the TGPSC syllabus continues with “Techniques for conducting regional surveys, cluster and factor analysis, input-output techniques. Spatial Analysis- Understanding structure of urban areas, density patterns, forces of concentration and dispersal, spatial distribution analysis, land use survey. Methods of data collection, sampling size and methods.” This lesson covers the analytical methods that turn raw regional data into planning intelligence. Surveys collect the data; cluster/factor analysis groups and reduces; I-O models trace the economy; spatial analysis maps patterns. The Town Planning Assistant uses these techniques when preparing Master Plans, Comprehensive Mobility Plans, district planning documents, and project DPRs. The exam tests survey type identification, cluster/factor analysis logic, I-O table interpretation, and spatial analysis concepts (density, distribution, land use).
C. Mechanism in Words
-
Regional surveys are the systematic collection of data about a region’s physical, economic, social, and infrastructure characteristics — the empirical foundation for plan preparation. The classic framework traces to Patrick Geddes (1854–1932), the Scottish planner-sociologist who developed the “Valley Section” (a cross-section from mountains to sea showing how human occupations and settlements relate to natural zones) and the “Survey–Analysis–Plan” sequence — survey first (collect the facts), then analyse (interpret the facts), then plan (propose interventions). Geddes’ sequence remains the global standard for plan preparation; the URDPFI 2015 guidelines operationalise it for Indian conditions. Regional survey types include: physical surveys (topography, geology, soils, climate, water); land use surveys (existing use of every parcel); demographic surveys (population, age-sex, households, migration); economic surveys (employment, occupations, industries, trade); social surveys (caste, religion, education, health); infrastructure surveys (water supply, sewerage, drainage, transport, power); housing surveys (type, condition, tenure); and institutional surveys (ULBs, parastatals, line departments).
-
Patrick Geddes’ contribution goes beyond the Survey–Analysis–Plan sequence. He argued that planners must understand the “place-work-folk” triad — the physical setting (place), the economic activity (work), and the social/cultural life (folk) of a city. His “Cities in Evolution” (1915) proposed the “conurbation” concept — contiguous urban areas that grow together (Mumbai-Pune, Delhi-Gurgaon-Faridabad, Hyderabad Urban Agglomeration). Geddes also pioneered “conservative surgery” — minimal intervention in historic urban fabric, in contrast to the wholesale demolition of Haussmann’s Paris or Lutyen’s imperial New Delhi. His influence on Indian planning is direct — Geddes worked on several Indian town plans in the 1910s–20s, including Balrampur, Baroda, Calcutta, and others; many of his ideas echo in modern Master Plan practice.
-
Cluster analysis is a multivariate statistical technique that groups observations into clusters of high internal similarity. Given many observations (say, all 600+ districts of India) described by many variables (per-capita income, literacy, female WPR, urban share, infant mortality, etc.), cluster analysis groups districts that are similar across all variables simultaneously. The most common algorithms are hierarchical clustering (which builds a tree of nested clusters) and k-means clustering (which assigns each observation to one of k clusters to minimise within-cluster variance). Planners use cluster analysis to identify “types” of regions — backward districts, prosperous districts, transitional districts — for targeted policy intervention. The Aspirational Districts Programme (Lesson 10.2) is a form of manual clustering; statistical cluster analysis provides the formal underpinning.
-
Factor analysis is a multivariate technique that reduces many correlated variables to a smaller number of underlying “factors” or dimensions. For example, given 20 economic indicators for districts, factor analysis might identify 3 underlying factors: “industrial development,” “agricultural prosperity,” and “service-sector development.” Each district gets a score on each factor; planners can then map these factor scores spatially. Principal Component Analysis (PCA) is the most common factor-analysis technique. Factor analysis is the inverse of cluster analysis — factor analysis groups variables into factors; cluster analysis groups observations into clusters. Together they are powerful tools for understanding regional structure from large datasets.
-
Input-Output (I-O) analysis, developed by Wassily Leontief (Nobel Prize 1973), models the interdependencies between sectors of an economy. The I-O table is a square matrix: rows = producing sectors (agriculture, mining, manufacturing, services, etc.); columns = consuming sectors; cells = the value of goods/services sector i sells to sector j. The table also shows final demand (household consumption, investment, government, exports) and value added (wages, profits, taxes, imports). I-O analysis answers: how much does an increase in final demand for one sector (say, construction) require increased output in supplying sectors (steel, cement, labour, transport)? The “multiplier” is the total increase in economic output per unit increase in final demand — a key concept in regional economic impact analysis. Planners use I-O to assess the impact of large projects (a new steel plant, a metro, a port) on the regional economy.
-
Spatial analysis is the set of techniques for analysing the location, distribution, and pattern of geographic phenomena. Key concepts: density patterns — how population, employment, or land use varies across space (concentric decline from CBD, polycentric clusters, uniform sprawl); forces of concentration — agglomeration economies, transport nodes, natural resource concentrations (which attract economic activity to specific locations); forces of dispersal — high land prices, congestion, pollution (which push activity outward); spatial distribution — the statistical pattern of points or areas (random, clustered, dispersed); land use survey — the systematic mapping of existing land use across a planning area, classified by category (residential, commercial, industrial, public-semi-public, recreational, transport, agriculture, etc.). Modern spatial analysis uses GIS (covered in Lesson 15.1) to compute these patterns.
-
Data collection methods in planning divide into primary and secondary. Primary data — collected directly by the planner for a specific study — includes household surveys, traffic counts, field observations, remote sensing (satellite imagery, aerial photos), GIS mapping, and participatory appraisals (PRA — Participatory Rural Appraisal). Primary data is expensive but precise and tailored. Secondary data — already collected by someone else — includes Census data (population, age-sex, housing, economic activity), National Sample Survey (NSS), Sample Registration System (SRS, vital rates), NFHS, state and ULB records, satellite imagery archives, and published studies. Secondary data is cheap but may be outdated, aggregated at the wrong level, or unavailable for the specific question. Skilled planners combine both — secondary for context, primary for the specific planning question.
-
Sampling is the technique of surveying a subset (the sample) to make inferences about the whole (the population). Sample design has two dimensions: sample size and sampling method. Sample size depends on the desired precision (margin of error), confidence level (typically 95%), and population variability. For a typical household survey in a city of 100,000 households, a sample of 400 households gives a ±5% margin of error at 95% confidence — a common Indian planning sample. Larger samples (1,000–2,000) are used for higher precision. Sampling methods: simple random sampling — every household has an equal chance of selection (clean but inefficient); stratified random sampling — population divided into strata (e.g., income groups, wards) and samples drawn from each (ensures representation); cluster sampling — randomly selected clusters (e.g., a few polling booths) with all households in those clusters surveyed (cheaper); systematic sampling — every nth household on a list. Planners typically use stratified random sampling for household surveys.
D. Core Concept Explanations
C1. Regional survey types
| Type | What it covers | Primary tools |
|---|---|---|
| Physical | Topography, geology, soils, climate, water | Topo sheets; well logs; meteorological data |
| Land use | Existing use of every parcel | Field survey + satellite imagery + GIS |
| Demographic | Population, age-sex, households, migration | Census + field sample survey |
| Economic | Employment, industries, trade | Establishment survey + secondary data |
| Social | Caste, religion, education, health | Census + NSS + NFHS |
| Infrastructure | Water, sewerage, drainage, transport, power | ULB records + field inspection |
| Housing | Type, condition, tenure | Housing board / Census housing tables |
| Institutional | ULBs, parastatals, line departments | Org charts + jurisdiction maps |
C2. Cluster vs Factor analysis
| Dimension | Cluster analysis | Factor analysis |
|---|---|---|
| What it groups | Observations (districts, households) | Variables (income, literacy, mortality) |
| Output | Clusters of similar observations | Factors (underlying dimensions) |
| Use case | Identify “types” of regions for targeted policy | Reduce many correlated variables to fewer dimensions |
| Typical algorithm | k-means; hierarchical | Principal Component Analysis (PCA) |
C3. I-O table structure
| Sector 1 | Sector 2 | … | Final demand | Total output | |
|---|---|---|---|---|---|
| Sector 1 | x_11 | x_12 | … | d_1 | X_1 |
| Sector 2 | x_21 | x_22 | … | d_2 | X_2 |
| … | … | … | … | … | … |
| Value added | v_1 | v_2 | … | ||
| Total input | X_1 | X_2 | … |
The cell x_ij = value sector i sells to sector j. The multiplier = total output / final demand change.
C4. Spatial analysis concepts
| Concept | Definition | Example |
|---|---|---|
| Density pattern | Distribution of population/activity across space | Concentric decline from CBD |
| Concentration forces | What attracts activity to specific locations | Agglomeration; transport nodes |
| Dispersal forces | What pushes activity outward | High land prices; congestion |
| Spatial distribution | Statistical pattern of points | Clustered, random, dispersed |
| Land use survey | Mapping existing use of every parcel | URDPFI land use classification |
C5. Sampling methods
| Method | Description | Best for |
|---|---|---|
| Simple random | Equal probability for every unit | Homogeneous populations |
| Stratified random | Population split into strata; sample each | Heterogeneous populations (income, ward) |
| Cluster | Sample clusters (e.g., polling booths); survey all in cluster | Cost reduction; widespread populations |
| Systematic | Every nth unit from a list | Orderly lists (voter rolls) |
E. Worked Numericals and Parameter Tables
E1. Sample size computation (simplified)
For a 95% confidence level, 5% margin of error, in a large population, the typical sample size is ~400 (using the simplified formula n = 1.96² / (2 × 0.05)² ≈ 384). For a margin of error of 3%: n ≈ 1.96² / (2 × 0.03)² ≈ 1,068. Larger samples → tighter precision.
E2. I-O multiplier
A region’s I-O table shows that for every ₹1 of construction output, the supplying sectors (steel, cement, labour, transport) produce an additional ₹0.60 of output. The total multiplier = 1 + 0.6 + 0.6 × 0.6 (next-round effects) + … ≈ 1 / (1 − 0.6) = 2.5 (simplified). A ₹100 crore construction project generates total economic output of ₹250 crore in the region.
E3. Land use composition
A city’s planning area is 10,000 hectares. Existing land use survey:
- Residential: 4,500 ha (45%)
- Commercial: 500 ha (5%)
- Industrial: 800 ha (8%)
- Public-semi-public: 1,000 ha (10%)
- Recreational: 700 ha (7%)
- Transport: 1,500 ha (15%)
- Agriculture/water body/other: 1,000 ha (10%)
URDPFI norms suggest 40–55% residential, 4–6% commercial, 8–12% industrial, 12–18% transport — so this city is close to norms but slightly low on recreational (URDPFI suggests 8–12%).
E4. Density gradient computation
A metro has the following gross density (pph) by distance from CBD:
- 0–5 km: 350
- 5–10 km: 250
- 10–15 km: 180
- 15–20 km: 120
- 20–25 km: 70
This is a negative density gradient — density declines with distance from CBD. The Clark (1951) exponential model: D(r) = D_0 × e^(−br), where D_0 is CBD density, b is the gradient. Fitting the data, b ≈ 0.06 per km.
F. Design Criteria
| Parameter | Standard / Typical value | Source |
|---|---|---|
| Geddes’ “Cities in Evolution” | 1915 | Geddes |
| Leontief Nobel Prize | 1973 (for I-O analysis) | Nobel committee |
| URDPFI residential land use | 40–55% of planning area | URDPFI 2015 |
| URDPFI commercial | 4–6% | URDPFI 2015 |
| URDPFI industrial | 8–12% | URDPFI 2015 |
| URDPFI transport | 12–18% | URDPFI 2015 |
| URDPFI recreational | 8–12% | URDPFI 2015 |
| Typical household survey sample (city) | ~400 (for ±5% at 95% confidence) | Statistical convention |
G. Application Zones
- Master Plan preparation — survey, analysis, plan per Geddes’ sequence.
- Regional planning — cluster/factor analysis of districts; I-O impact of major projects.
- Project DPRs — land use survey, demographic projection, infrastructure demand.
- Policy targeting — cluster analysis for identifying backward districts/areas.
- Housing surveys — household sampling for affordability, condition, tenure.
H. Common Confusions
| Confusion | Reality |
|---|---|
| “Cluster analysis groups variables.” | No — cluster analysis groups observations (e.g., districts). Factor analysis groups variables. |
| “I-O analysis was developed by Keynes.” | No — by Wassily Leontief (Nobel 1973). |
| “Geddes was an architect.” | Not primarily — Geddes was a botanist-sociologist-planner; his contribution is the Survey–Analysis–Plan sequence, not architectural design. |
| “Stratified sampling and cluster sampling are the same.” | No — stratified splits population into strata and samples each; cluster samples whole clusters. |
| “Higher sample size always gives better results.” | True for precision, but with diminishing returns; depends on sampling method and quality. |
| “Land use survey is optional for a Master Plan.” | False — land use survey is a mandatory step in plan preparation per URDPFI. |
| “Primary data is always better than secondary.” | Not always — primary is more current and tailored but expensive; secondary is cheaper and often authoritative (e.g., Census). |
I. Compare & Contrast
I1. Cluster vs Factor analysis
| Dimension | Cluster analysis | Factor analysis |
|---|---|---|
| Unit grouped | Observations | Variables |
| Output | Clusters | Factors (underlying dimensions) |
| Example | Group districts by similarity | Reduce 20 indicators to 3 factors |
I2. Primary vs Secondary data
| Dimension | Primary data | Secondary data |
|---|---|---|
| Source | Directly collected | Already published |
| Cost | High | Low |
| Precision | High (tailored) | Varies |
| Timeliness | Current | May be outdated |
| Examples | Household surveys; traffic counts; remote sensing | Census; NSS; NFHS; ULB records |
J. Memory Hooks
- “Geddes: Valley Section + Survey-Analysis-Plan + Place-Work-Folk” — three foundational contributions.
- “Cluster groups observations; Factor groups variables” — the key distinction.
- “Leontief → I-O → multiplier” — the Nobel laureate + technique + key output.
- “D(r) = D_0 × e^(−br)” — the negative density gradient (Clark model).
- “Sample size: 400 = ±5% at 95%; 1,000 = ±3%” — typical planning samples.
- “URDPFI land use: R-C-I-P-R-T = 45-5-10-10-10-15%” — approximate planning norms.
- “Sampling: Simple-Stratified-Cluster-Systematic” — four methods.
K. Revision Ladder
| Order | Item | Time |
|---|---|---|
| 1 | Memorise Geddes’ three contributions | 20 min |
| 2 | Memorise regional survey types | 30 min |
| 3 | Memorise cluster vs factor analysis distinction | 20 min |
| 4 | Memorise I-O analysis structure and multiplier concept | 30 min |
| 5 | Memorise spatial analysis concepts (density, forces, distribution) | 30 min |
| 6 | Memorise URDPFI land use norms | 30 min |
| 7 | Memorise primary vs secondary data with examples | 20 min |
| 8 | Memorise the 4 sampling methods | 20 min |
| 9 | Practise sample size and I-O multiplier arithmetic | 30 min |
L. Exam Traps
| Trap | Correct response |
|---|---|
| Question pairs cluster analysis with grouping variables. | False — cluster groups observations; factor groups variables. |
| Question pairs I-O analysis with Keynes. | False — Leontief (Nobel 1973). |
| Question pairs Geddes with architecture. | False — Geddes was a botanist-sociologist-planner; his contribution is the Survey-Analysis-Plan sequence. |
| Question pairs stratified with cluster sampling. | They are different — stratified splits into strata and samples each; cluster samples whole clusters. |
| Question lists 400-sample as ±3% margin. | False — 400 gives ±5% (at 95% confidence); ±3% needs ~1,000. |
| Question pairs land use survey as optional for Master Plan. | False — land use survey is mandatory per URDPFI. |
M. Answer-Writing Cues
- For Geddes questions, give year + work + three contributions: “Patrick Geddes, in ‘Cities in Evolution’ (1915), contributed the Valley Section, the Survey-Analysis-Plan sequence, and the Place-Work-Folk triad — foundational ideas that still guide plan preparation.”
- For I-O questions, give developer + structure + multiplier concept.
- For sampling questions, give method + sample size formula + worked example.
- For data collection, give primary vs secondary with examples.
N. PYQ Integration
Pattern questions only:
Pattern question 1 — Patrick Geddes
Q. The “Survey–Analysis–Plan” sequence, foundational to modern plan preparation, was developed by:
– (A) Ebenezer Howard
– (B) Patrick Geddes ✓
– (C) Le Corbusier
– (D) Clarence Stein
Ans: (B). Geddes also developed the Valley Section and the Place-Work-Folk triad.
Pattern question 2 — Cluster vs Factor
Q. Cluster analysis is a multivariate statistical technique that:
– (A) Groups observations into clusters of high internal similarity ✓
– (B) Reduces many variables to fewer underlying factors
– (C) Models interdependencies between economic sectors
– (D) Fits an S-curve to historical data
Ans: (A). Option (B) is factor analysis; (C) is I-O; (D) is logistic projection.
Pattern question 3 — I-O analysis
Q. Input-Output (I-O) analysis, developed by Wassily Leontief, is primarily used for:
– (A) Population projection
– (B) Modelling interdependencies between economic sectors ✓
– (C) Cluster grouping of districts
– (D) Sampling design
Ans: (B). I-O models the multiplier effects across sectors.
Pattern question 4 — MSQ
Q. Which of the following are standard regional survey types?
– (A) Physical survey ✓
– (B) Land use survey ✓
– (C) Demographic survey ✓
– (D) Astronomical survey
Ans: (A), (B), (C). Astronomical surveys are not planning-related.
Pattern question 5 — Numerical
An I-O table shows that for every ₹1 increase in construction output, supplying sectors produce an additional ₹0.50 of output. The simplified multiplier is approximately:
– (A) 1.5
– (B) 2.0 ✓
– (C) 2.5
– (D) 5.0
Ans: (B). 1 / (1 − 0.5) = 2.0.
O. Mini-Check — Lesson 12.2
- List Patrick Geddes’ three principal contributions to planning.
- State the Survey-Analysis-Plan sequence and its originator.
- Distinguish cluster analysis from factor analysis.
- Who developed I-O analysis, and what does it model?
- State the principal concepts in spatial analysis (any four).
- State the difference between primary and secondary data, with one example each.
- List the four standard sampling methods.
- What sample size gives ±5% margin at 95% confidence in a typical household survey?
- State the URDPFI norm for residential land use as a share of planning area.
- Distinguish stratified random sampling from cluster sampling.
Answers:
1. Valley Section (cross-section showing human occupations by natural zone); Survey-Analysis-Plan sequence (collect facts, interpret, propose); Place-Work-Folk triad (physical setting, economic activity, social life).
2. Survey first, then analyse, then plan — developed by Patrick Geddes (in “Cities in Evolution,” 1915).
3. Cluster analysis groups observations (e.g., districts) into similar clusters. Factor analysis reduces many correlated variables to fewer underlying factors.
4. Wassily Leontief (Nobel Prize 1973). I-O models the interdependencies between economic sectors — how increased final demand in one sector requires increased output in supplying sectors, producing multiplier effects.
5. Density patterns; forces of concentration (agglomeration, nodes); forces of dispersal (high land prices, congestion); spatial distribution (random, clustered, dispersed); land use survey. Any four.
6. Primary data = directly collected for the study (e.g., household survey, traffic count, satellite image); secondary data = already published (e.g., Census, NSS, ULB records). Primary is current but expensive; secondary is cheap but may be outdated.
7. Simple random, stratified random, cluster, systematic.
8. ~400 households (for a large population).
9. 40–55% of planning area.
10. Stratified splits the population into strata (e.g., wards, income groups) and samples from each; cluster samples whole clusters (e.g., randomly selected polling booths) and surveys all households in those clusters.
Module 12 complete. Next: Module 13 — Urban Design & Conservation (Paper II §13, 1 lesson covering Lynch, urban form, and conservation framework).