David Britain – tezės

Regional dialectology has, for over 150 years, provided an invaluable historical baseline for the study of linguistic variation, from, to use British examples, Bonaparte’s and Ellis’s nineteenth-century surveys of „uncultivated peasants” through to the monumental Survey of English Dialects (SED, 1962–71). These foundational projects achieved impressive regional coverage across hundreds of sites, but their methodological compromises are equally instructive: restricted demographic sampling to Non-mobile, Older, Rural Males (NORMs); on-the-spot rather than recorded transcription; and data-collection timescales stretching to decades. Labovian sociolinguistics from the 1960s redirected attention toward socially stratified variation within single communities, improving representativeness, data capture, and speed—but at the cost of multilocality coverage, leaving researchers today still reliant on SED data drawn from nineteenth-century-born speakers.

Computational reanalyses of legacy SED data (Viereck 1986; Shackleton 2007) offered early quantitative inspections of historical corpora, while Web 2.0 has transformed data collection itself: internet quizzes, social media scraping, and smartphone applications such as the English Dialects App and the BBC Future „secret accent” survey have generated datasets of unprecedented scale at a fraction of the cost and time of traditional fieldwork, while beginning to incorporate social metadata largely absent from earlier atlases. These crowdsourced methods nonetheless remain self-report, sampling-biased, and text- rather than speech-based, leaving a persistent gap: no nationwide, speech-based survey of England has combined sociolinguistic sensitivity with genuine regional coverage.

The talk then turns to CURLEW (Census of Urban and Regional Language in England and Wales), a new four-year project designed to close that gap for England and Wales. Modelled on Leemann’s Swiss Dialäktatlas, CURLEW will survey over 140 sites across England and Wales, chosen to maximise dialect diversity and population spread, and will recruit a demographically broader range of speakers than SED’s NORM-only sampling. Each participant completes some 300 target-variable elicitation items spanning phonetics, lexis, morphology, syntax, and discourse, plus a read-speech passage and spontaneous conversation.

Data collection is conducted remotely via smartphone app with Zoom-supervised fieldwork, while AI tools automate the processing that made earlier surveys so time-intensive. My talk reflects candidly on the compromises this entails, arguing that CURLEW offers a template for a technologically enabled successor to the Survey of English Dialects, a century on.