LESSON 12.2 — Regional Surveys, Cluster/Factor Analysis, I-O Techniques & Spatial Analysis

A. Standard Map

Topic Governing Source Exam Focus
Regional surveys Multiple types for plan preparation Survey types + purpose
Patrick Geddes “Valley Section” + survey–analysis–plan Foundational logic
Cluster analysis Statistical grouping of similar observations Method + use
Factor analysis Reduce many variables to fewer dimensions Method + use
Input-Output (I-O) techniques Sector × sector economic transactions matrix Method + use
Spatial analysis Density patterns, distribution, land use Concepts + measures
Data collection methods Primary (surveys, RS, GIS) vs secondary (Census) Sources
Sampling Size and method Sample arithmetic

B. Why It’s Used

Paper II §12 of the TGPSC syllabus continues with “Techniques for conducting regional surveys, cluster and factor analysis, input-output techniques. Spatial Analysis- Understanding structure of urban areas, density patterns, forces of concentration and dispersal, spatial distribution analysis, land use survey. Methods of data collection, sampling size and methods.” This lesson covers the analytical methods that turn raw regional data into planning intelligence. Surveys collect the data; cluster/factor analysis groups and reduces; I-O models trace the economy; spatial analysis maps patterns. The Town Planning Assistant uses these techniques when preparing Master Plans, Comprehensive Mobility Plans, district planning documents, and project DPRs. The exam tests survey type identification, cluster/factor analysis logic, I-O table interpretation, and spatial analysis concepts (density, distribution, land use).


C. Mechanism in Words

  1. Regional surveys are the systematic collection of data about a region’s physical, economic, social, and infrastructure characteristics — the empirical foundation for plan preparation. The classic framework traces to Patrick Geddes (1854–1932), the Scottish planner-sociologist who developed the “Valley Section” (a cross-section from mountains to sea showing how human occupations and settlements relate to natural zones) and the “Survey–Analysis–Plan” sequence — survey first (collect the facts), then analyse (interpret the facts), then plan (propose interventions). Geddes’ sequence remains the global standard for plan preparation; the URDPFI 2015 guidelines operationalise it for Indian conditions. Regional survey types include: physical surveys (topography, geology, soils, climate, water); land use surveys (existing use of every parcel); demographic surveys (population, age-sex, households, migration); economic surveys (employment, occupations, industries, trade); social surveys (caste, religion, education, health); infrastructure surveys (water supply, sewerage, drainage, transport, power); housing surveys (type, condition, tenure); and institutional surveys (ULBs, parastatals, line departments).

  2. Patrick Geddes’ contribution goes beyond the Survey–Analysis–Plan sequence. He argued that planners must understand the “place-work-folk” triad — the physical setting (place), the economic activity (work), and the social/cultural life (folk) of a city. His “Cities in Evolution” (1915) proposed the “conurbation” concept — contiguous urban areas that grow together (Mumbai-Pune, Delhi-Gurgaon-Faridabad, Hyderabad Urban Agglomeration). Geddes also pioneered “conservative surgery” — minimal intervention in historic urban fabric, in contrast to the wholesale demolition of Haussmann’s Paris or Lutyen’s imperial New Delhi. His influence on Indian planning is direct — Geddes worked on several Indian town plans in the 1910s–20s, including Balrampur, Baroda, Calcutta, and others; many of his ideas echo in modern Master Plan practice.

  3. Cluster analysis is a multivariate statistical technique that groups observations into clusters of high internal similarity. Given many observations (say, all 600+ districts of India) described by many variables (per-capita income, literacy, female WPR, urban share, infant mortality, etc.), cluster analysis groups districts that are similar across all variables simultaneously. The most common algorithms are hierarchical clustering (which builds a tree of nested clusters) and k-means clustering (which assigns each observation to one of k clusters to minimise within-cluster variance). Planners use cluster analysis to identify “types” of regions — backward districts, prosperous districts, transitional districts — for targeted policy intervention. The Aspirational Districts Programme (Lesson 10.2) is a form of manual clustering; statistical cluster analysis provides the formal underpinning.

  4. Factor analysis is a multivariate technique that reduces many correlated variables to a smaller number of underlying “factors” or dimensions. For example, given 20 economic indicators for districts, factor analysis might identify 3 underlying factors: “industrial development,” “agricultural prosperity,” and “service-sector development.” Each district gets a score on each factor; planners can then map these factor scores spatially. Principal Component Analysis (PCA) is the most common factor-analysis technique. Factor analysis is the inverse of cluster analysis — factor analysis groups variables into factors; cluster analysis groups observations into clusters. Together they are powerful tools for understanding regional structure from large datasets.

  5. Input-Output (I-O) analysis, developed by Wassily Leontief (Nobel Prize 1973), models the interdependencies between sectors of an economy. The I-O table is a square matrix: rows = producing sectors (agriculture, mining, manufacturing, services, etc.); columns = consuming sectors; cells = the value of goods/services sector i sells to sector j. The table also shows final demand (household consumption, investment, government, exports) and value added (wages, profits, taxes, imports). I-O analysis answers: how much does an increase in final demand for one sector (say, construction) require increased output in supplying sectors (steel, cement, labour, transport)? The “multiplier” is the total increase in economic output per unit increase in final demand — a key concept in regional economic impact analysis. Planners use I-O to assess the impact of large projects (a new steel plant, a metro, a port) on the regional economy.

  6. Spatial analysis is the set of techniques for analysing the location, distribution, and pattern of geographic phenomena. Key concepts: density patterns — how population, employment, or land use varies across space (concentric decline from CBD, polycentric clusters, uniform sprawl); forces of concentration — agglomeration economies, transport nodes, natural resource concentrations (which attract economic activity to specific locations); forces of dispersal — high land prices, congestion, pollution (which push activity outward); spatial distribution — the statistical pattern of points or areas (random, clustered, dispersed); land use survey — the systematic mapping of existing land use across a planning area, classified by category (residential, commercial, industrial, public-semi-public, recreational, transport, agriculture, etc.). Modern spatial analysis uses GIS (covered in Lesson 15.1) to compute these patterns.

  7. Data collection methods in planning divide into primary and secondary. Primary data — collected directly by the planner for a specific study — includes household surveys, traffic counts, field observations, remote sensing (satellite imagery, aerial photos), GIS mapping, and participatory appraisals (PRA — Participatory Rural Appraisal). Primary data is expensive but precise and tailored. Secondary data — already collected by someone else — includes Census data (population, age-sex, housing, economic activity), National Sample Survey (NSS), Sample Registration System (SRS, vital rates), NFHS, state and ULB records, satellite imagery archives, and published studies. Secondary data is cheap but may be outdated, aggregated at the wrong level, or unavailable for the specific question. Skilled planners combine both — secondary for context, primary for the specific planning question.

  8. Sampling is the technique of surveying a subset (the sample) to make inferences about the whole (the population). Sample design has two dimensions: sample size and sampling method. Sample size depends on the desired precision (margin of error), confidence level (typically 95%), and population variability. For a typical household survey in a city of 100,000 households, a sample of 400 households gives a ±5% margin of error at 95% confidence — a common Indian planning sample. Larger samples (1,000–2,000) are used for higher precision. Sampling methods: simple random sampling — every household has an equal chance of selection (clean but inefficient); stratified random sampling — population divided into strata (e.g., income groups, wards) and samples drawn from each (ensures representation); cluster sampling — randomly selected clusters (e.g., a few polling booths) with all households in those clusters surveyed (cheaper); systematic sampling — every nth household on a list. Planners typically use stratified random sampling for household surveys.


D. Core Concept Explanations

C1. Regional survey types

Type What it covers Primary tools
Physical Topography, geology, soils, climate, water Topo sheets; well logs; meteorological data
Land use Existing use of every parcel Field survey + satellite imagery + GIS
Demographic Population, age-sex, households, migration Census + field sample survey
Economic Employment, industries, trade Establishment survey + secondary data
Social Caste, religion, education, health Census + NSS + NFHS
Infrastructure Water, sewerage, drainage, transport, power ULB records + field inspection
Housing Type, condition, tenure Housing board / Census housing tables
Institutional ULBs, parastatals, line departments Org charts + jurisdiction maps

C2. Cluster vs Factor analysis

Dimension Cluster analysis Factor analysis
What it groups Observations (districts, households) Variables (income, literacy, mortality)
Output Clusters of similar observations Factors (underlying dimensions)
Use case Identify “types” of regions for targeted policy Reduce many correlated variables to fewer dimensions
Typical algorithm k-means; hierarchical Principal Component Analysis (PCA)

C3. I-O table structure

Sector 1 Sector 2 Final demand Total output
Sector 1 x_11 x_12 d_1 X_1
Sector 2 x_21 x_22 d_2 X_2
Value added v_1 v_2
Total input X_1 X_2

The cell x_ij = value sector i sells to sector j. The multiplier = total output / final demand change.

C4. Spatial analysis concepts

Concept Definition Example
Density pattern Distribution of population/activity across space Concentric decline from CBD
Concentration forces What attracts activity to specific locations Agglomeration; transport nodes
Dispersal forces What pushes activity outward High land prices; congestion
Spatial distribution Statistical pattern of points Clustered, random, dispersed
Land use survey Mapping existing use of every parcel URDPFI land use classification

C5. Sampling methods

Method Description Best for
Simple random Equal probability for every unit Homogeneous populations
Stratified random Population split into strata; sample each Heterogeneous populations (income, ward)
Cluster Sample clusters (e.g., polling booths); survey all in cluster Cost reduction; widespread populations
Systematic Every nth unit from a list Orderly lists (voter rolls)

E. Worked Numericals and Parameter Tables

E1. Sample size computation (simplified)

For a 95% confidence level, 5% margin of error, in a large population, the typical sample size is ~400 (using the simplified formula n = 1.96² / (2 × 0.05)² ≈ 384). For a margin of error of 3%: n ≈ 1.96² / (2 × 0.03)² ≈ 1,068. Larger samples → tighter precision.

E2. I-O multiplier

A region’s I-O table shows that for every ₹1 of construction output, the supplying sectors (steel, cement, labour, transport) produce an additional ₹0.60 of output. The total multiplier = 1 + 0.6 + 0.6 × 0.6 (next-round effects) + … ≈ 1 / (1 − 0.6) = 2.5 (simplified). A ₹100 crore construction project generates total economic output of ₹250 crore in the region.

E3. Land use composition

A city’s planning area is 10,000 hectares. Existing land use survey:

  • Residential: 4,500 ha (45%)
  • Commercial: 500 ha (5%)
  • Industrial: 800 ha (8%)
  • Public-semi-public: 1,000 ha (10%)
  • Recreational: 700 ha (7%)
  • Transport: 1,500 ha (15%)
  • Agriculture/water body/other: 1,000 ha (10%)

URDPFI norms suggest 40–55% residential, 4–6% commercial, 8–12% industrial, 12–18% transport — so this city is close to norms but slightly low on recreational (URDPFI suggests 8–12%).

E4. Density gradient computation

A metro has the following gross density (pph) by distance from CBD:

  • 0–5 km: 350
  • 5–10 km: 250
  • 10–15 km: 180
  • 15–20 km: 120
  • 20–25 km: 70

This is a negative density gradient — density declines with distance from CBD. The Clark (1951) exponential model: D(r) = D_0 × e^(−br), where D_0 is CBD density, b is the gradient. Fitting the data, b ≈ 0.06 per km.


F. Design Criteria

Parameter Standard / Typical value Source
Geddes’ “Cities in Evolution” 1915 Geddes
Leontief Nobel Prize 1973 (for I-O analysis) Nobel committee
URDPFI residential land use 40–55% of planning area URDPFI 2015
URDPFI commercial 4–6% URDPFI 2015
URDPFI industrial 8–12% URDPFI 2015
URDPFI transport 12–18% URDPFI 2015
URDPFI recreational 8–12% URDPFI 2015
Typical household survey sample (city) ~400 (for ±5% at 95% confidence) Statistical convention

G. Application Zones

  1. Master Plan preparation — survey, analysis, plan per Geddes’ sequence.
  2. Regional planning — cluster/factor analysis of districts; I-O impact of major projects.
  3. Project DPRs — land use survey, demographic projection, infrastructure demand.
  4. Policy targeting — cluster analysis for identifying backward districts/areas.
  5. Housing surveys — household sampling for affordability, condition, tenure.

H. Common Confusions

Confusion Reality
“Cluster analysis groups variables.” No — cluster analysis groups observations (e.g., districts). Factor analysis groups variables.
“I-O analysis was developed by Keynes.” No — by Wassily Leontief (Nobel 1973).
“Geddes was an architect.” Not primarily — Geddes was a botanist-sociologist-planner; his contribution is the Survey–Analysis–Plan sequence, not architectural design.
“Stratified sampling and cluster sampling are the same.” No — stratified splits population into strata and samples each; cluster samples whole clusters.
“Higher sample size always gives better results.” True for precision, but with diminishing returns; depends on sampling method and quality.
“Land use survey is optional for a Master Plan.” False — land use survey is a mandatory step in plan preparation per URDPFI.
“Primary data is always better than secondary.” Not always — primary is more current and tailored but expensive; secondary is cheaper and often authoritative (e.g., Census).

I. Compare & Contrast

I1. Cluster vs Factor analysis

Dimension Cluster analysis Factor analysis
Unit grouped Observations Variables
Output Clusters Factors (underlying dimensions)
Example Group districts by similarity Reduce 20 indicators to 3 factors

I2. Primary vs Secondary data

Dimension Primary data Secondary data
Source Directly collected Already published
Cost High Low
Precision High (tailored) Varies
Timeliness Current May be outdated
Examples Household surveys; traffic counts; remote sensing Census; NSS; NFHS; ULB records

J. Memory Hooks

  • “Geddes: Valley Section + Survey-Analysis-Plan + Place-Work-Folk” — three foundational contributions.
  • “Cluster groups observations; Factor groups variables” — the key distinction.
  • “Leontief → I-O → multiplier” — the Nobel laureate + technique + key output.
  • “D(r) = D_0 × e^(−br)” — the negative density gradient (Clark model).
  • “Sample size: 400 = ±5% at 95%; 1,000 = ±3%” — typical planning samples.
  • “URDPFI land use: R-C-I-P-R-T = 45-5-10-10-10-15%” — approximate planning norms.
  • “Sampling: Simple-Stratified-Cluster-Systematic” — four methods.

K. Revision Ladder

Order Item Time
1 Memorise Geddes’ three contributions 20 min
2 Memorise regional survey types 30 min
3 Memorise cluster vs factor analysis distinction 20 min
4 Memorise I-O analysis structure and multiplier concept 30 min
5 Memorise spatial analysis concepts (density, forces, distribution) 30 min
6 Memorise URDPFI land use norms 30 min
7 Memorise primary vs secondary data with examples 20 min
8 Memorise the 4 sampling methods 20 min
9 Practise sample size and I-O multiplier arithmetic 30 min

L. Exam Traps

Trap Correct response
Question pairs cluster analysis with grouping variables. False — cluster groups observations; factor groups variables.
Question pairs I-O analysis with Keynes. False — Leontief (Nobel 1973).
Question pairs Geddes with architecture. False — Geddes was a botanist-sociologist-planner; his contribution is the Survey-Analysis-Plan sequence.
Question pairs stratified with cluster sampling. They are different — stratified splits into strata and samples each; cluster samples whole clusters.
Question lists 400-sample as ±3% margin. False — 400 gives ±5% (at 95% confidence); ±3% needs ~1,000.
Question pairs land use survey as optional for Master Plan. False — land use survey is mandatory per URDPFI.

M. Answer-Writing Cues

  • For Geddes questions, give year + work + three contributions: “Patrick Geddes, in ‘Cities in Evolution’ (1915), contributed the Valley Section, the Survey-Analysis-Plan sequence, and the Place-Work-Folk triad — foundational ideas that still guide plan preparation.”
  • For I-O questions, give developer + structure + multiplier concept.
  • For sampling questions, give method + sample size formula + worked example.
  • For data collection, give primary vs secondary with examples.

N. PYQ Integration

Pattern questions only:

Pattern question 1 — Patrick Geddes

Q. The “Survey–Analysis–Plan” sequence, foundational to modern plan preparation, was developed by:
– (A) Ebenezer Howard
– (B) Patrick Geddes ✓
– (C) Le Corbusier
– (D) Clarence Stein

Ans: (B). Geddes also developed the Valley Section and the Place-Work-Folk triad.

Pattern question 2 — Cluster vs Factor

Q. Cluster analysis is a multivariate statistical technique that:
– (A) Groups observations into clusters of high internal similarity ✓
– (B) Reduces many variables to fewer underlying factors
– (C) Models interdependencies between economic sectors
– (D) Fits an S-curve to historical data

Ans: (A). Option (B) is factor analysis; (C) is I-O; (D) is logistic projection.

Pattern question 3 — I-O analysis

Q. Input-Output (I-O) analysis, developed by Wassily Leontief, is primarily used for:
– (A) Population projection
– (B) Modelling interdependencies between economic sectors ✓
– (C) Cluster grouping of districts
– (D) Sampling design

Ans: (B). I-O models the multiplier effects across sectors.

Pattern question 4 — MSQ

Q. Which of the following are standard regional survey types?
– (A) Physical survey ✓
– (B) Land use survey ✓
– (C) Demographic survey ✓
– (D) Astronomical survey

Ans: (A), (B), (C). Astronomical surveys are not planning-related.

Pattern question 5 — Numerical

An I-O table shows that for every ₹1 increase in construction output, supplying sectors produce an additional ₹0.50 of output. The simplified multiplier is approximately:
– (A) 1.5
– (B) 2.0 ✓
– (C) 2.5
– (D) 5.0

Ans: (B). 1 / (1 − 0.5) = 2.0.


O. Mini-Check — Lesson 12.2

  1. List Patrick Geddes’ three principal contributions to planning.
  2. State the Survey-Analysis-Plan sequence and its originator.
  3. Distinguish cluster analysis from factor analysis.
  4. Who developed I-O analysis, and what does it model?
  5. State the principal concepts in spatial analysis (any four).
  6. State the difference between primary and secondary data, with one example each.
  7. List the four standard sampling methods.
  8. What sample size gives ±5% margin at 95% confidence in a typical household survey?
  9. State the URDPFI norm for residential land use as a share of planning area.
  10. Distinguish stratified random sampling from cluster sampling.

Answers:
1. Valley Section (cross-section showing human occupations by natural zone); Survey-Analysis-Plan sequence (collect facts, interpret, propose); Place-Work-Folk triad (physical setting, economic activity, social life).
2. Survey first, then analyse, then plan — developed by Patrick Geddes (in “Cities in Evolution,” 1915).
3. Cluster analysis groups observations (e.g., districts) into similar clusters. Factor analysis reduces many correlated variables to fewer underlying factors.
4. Wassily Leontief (Nobel Prize 1973). I-O models the interdependencies between economic sectors — how increased final demand in one sector requires increased output in supplying sectors, producing multiplier effects.
5. Density patterns; forces of concentration (agglomeration, nodes); forces of dispersal (high land prices, congestion); spatial distribution (random, clustered, dispersed); land use survey. Any four.
6. Primary data = directly collected for the study (e.g., household survey, traffic count, satellite image); secondary data = already published (e.g., Census, NSS, ULB records). Primary is current but expensive; secondary is cheap but may be outdated.
7. Simple random, stratified random, cluster, systematic.
8. ~400 households (for a large population).
9. 40–55% of planning area.
10. Stratified splits the population into strata (e.g., wards, income groups) and samples from each; cluster samples whole clusters (e.g., randomly selected polling booths) and surveys all households in those clusters.


Module 12 complete. Next: Module 13 — Urban Design & Conservation (Paper II §13, 1 lesson covering Lynch, urban form, and conservation framework).