Skip to main content
Destination Depth Scoring

Destination Depth Scoring Decisions Under Time Pressure

Rankings flatten a city into a single number. That number might tell you how many people visited or how many hotels got booked, but it won't tell you whether a neighborhood feels alive at 9pm or whether the best taco stand is hidden behind a laundromat. Depth benchmarking is a different game. It measures layers—cultural, historical, sensory—so you can compare cities the way a curious traveler would, not the way a spreadsheet does. When teams treat this step as optional, the rework loop usually starts within one sprint because the baseline checklist never got logged, and reviewers spot the gap before anyone retests the failure mode in the field. Fix this part first. Watershed crews who keep phenology notes beside camera-trap cards treat absence as a process signal, not a missing checkbox, and that habit alone keeps seasonal reports from reading like cloned templates under review. Varroa nectar drifts sideways.

Rankings flatten a city into a single number. That number might tell you how many people visited or how many hotels got booked, but it won't tell you whether a neighborhood feels alive at 9pm or whether the best taco stand is hidden behind a laundromat. Depth benchmarking is a different game. It measures layers—cultural, historical, sensory—so you can compare cities the way a curious traveler would, not the way a spreadsheet does. When teams treat this step as optional, the rework loop usually starts within one sprint because the baseline checklist never got logged, and reviewers spot the gap before anyone retests the failure mode in the field. Fix this part first. Watershed crews who keep phenology notes beside camera-trap cards treat absence as a process signal, not a missing checkbox, and that habit alone keeps seasonal reports from reading like cloned templates under review. Varroa nectar drifts sideways. This field guide comes from real work: destination marketing teams, independent travel writers, and city planners who need to explain why one place feels richer than another. The tools aren't proprietary. The patterns are earned through trial and error. You'll find the pitfalls, the maintenance costs, and the moments when you should just quit and use a ranking instead.

Where Depth Benchmarks Actually Show Up in Travel Work

Ask anyone who runs a destination account, and they'll admit it: the official tourism stats never match what they actually sell. You pitch a city on its food scene, but the visitor bureau gives you hotel occupancy rates. You build a campaign around street art, and the only numbers available are airport arrivals. So teams start inventing their own measures. I have sat in a cramped conference room where three people argued over how to score “walkability at 9pm”—that's a depth benchmark forming in the wild, whether anyone calls it that or not.

The usual workaround is a spreadsheet with weighted columns: number of independent shops per block, average length of a café conversation, how many locals say hello before you reach the second cross-street. None of this appears in any official report. But it drives the brochure images, the “hidden gems” listicles, and the paid partnerships with local guides. The catch is that these scores stay private because they feel unscientific. Teams fear the judgment of a finance director who wants a clean ROI chart. So the depth work happens in shadow, refined by trial and error, and rarely shared beyond a Slack channel.

Worth flagging—this secrecy is a mistake. When a score stays hidden, it never gets tested against reality. I have seen a city team rank neighborhoods by “authentic charm” only to discover their top choice had been gentrified for a decade. The benchmark worked, but nobody checked it against ground truth.

“You can rank a hundred cities by population, but that tells you nothing about whether you'd want to get lost in them.”

— former travel editor, on why she stopped publishing listicles

How independent writers use benchmarks without naming them

Freelance travel journalists do this constantly, though they'd never use the term. When a writer files a story about “three neighborhoods that feel completely different,” they're scoring depth by contrast. When they compare a tourist district to a residential pocket ten minutes away, they're measuring a layer that official guides skip. The trick is that writers trust lived signals over data tables.

One editor I worked with kept a personal ranking system for every city she visited: how many times she got lost and found something good, how often a local corrected her assumptions, whether the main square felt like a stage or a living room. She never wrote those scores down. But they determined which cities got a feature and which got a single paragraph. That's a depth benchmark in practice—informal, inconsistent, but brutally effective at separating places with surface appeal from places with real texture.

That sounds fine until you realize the gap between official stats and lived experience is where most travel writing fails. A city can have record visitor numbers and still feel hollow. A smaller town with half the tourists can leave you exhausted in the best way. Writers know this instinctively, which is why they develop private heuristics. The problem is that these heuristics stay individual. Nobody audits them, so they drift toward whatever the writer enjoyed last week. That's the trade-off: personal benchmarks feel honest but carry no discipline.

The gap between official stats and lived experience

Here's the concrete gap. A city's official profile lists museums, parks, and transit scores. Those are useful for a first glance, but they flatten everything into amenities. Depth requires asking different questions: does the bakery close at 2pm because it sells out, or because the owner just can't be bothered? Are the people at the bus stop talking to each other, or staring at their phones? Does the main square host a weekly market that locals actually shop at, or is it a photo backdrop?

Most teams never bridge this gap because the official data is free and instantly available. Pulling depth signals requires fieldwork—walking specific streets at three different times of day, interviewing shopkeepers, noticing which businesses survive a downtown rent hike. That's expensive work, and it doesn't scale across a hundred cities. So the default becomes shallow proxies: number of cafes per capita, density of art galleries, proximity to a river. These are not depth benchmarks. They're just rankings wearing a costume.

What usually breaks first is the moment a team realizes their depth score matches the tourist board's marketing material. That's a red flag. Real depth metrics should occasionally surprise you—a neighborhood you ignored that scores high, a famous district that falls flat. If your benchmark only confirms what you already think, it's measuring your bias, not the city.

Foundations People Keep Mixing Up

Depth vs. popularity: they split almost immediately. A city can be famous and shallow. Vegas has millions of visitors, a strip packed wall-to-wall with bodies—and almost no one knows where the water comes from, why the neon settled there, or what the valley looked like before the dams. Popularity measures how many people show up. Depth measures how many layers you can peel before hitting bedrock. The divergence isn't accidental; it's structural. Popularity rewards what's easy to consume, photograph, and check off. Depth rewards what resists consumption.

Depth vs. popularity: why they diverge

I have watched travel teams build a "must-see" list in an afternoon—ten attractions, all Instagram-verified, zero context. Then they wonder why the itinerary feels hollow by day three. The catch is that popularity data is cheap. You scrape it, sort it, and call it research. Depth data costs hours—walking, asking, failing to find the entrance, backtracking. That's why the shortcuts multiply.

Layers vs. categories: a key distinction

Categories are containers; layers are strata. A category says "this is food." A layer says "this is the fourth-generation family stall that survived the 2008 recession, the 2015 rent hike, and the pandemic, and the grandson just added a vegan option that the grandmother quietly disapproves of." Categories flatten. Layers stack. Most scoring systems stop at categories because they're tidy. Layers are messy—they have contradictions, dead ends, and competing narratives.

The trap I see repeatedly: teams build a taxonomy of venue types (museums, markets, parks) and call it depth. Wrong order. That's just a fancier list. Layers require temporal depth (how things changed), relational depth (who connects to whom), and experiential depth (what it actually feels like to be there). A market categorized as "food" tells you nothing. A market layered with its dawn rituals, the fishmonger's lineage, and the after-hours cleaning crew—that's depth.

Honestly — most travel posts skip this.

Most teams skip this distinction entirely. They grab categories because they're defensible—you can point to a spreadsheet and say "we covered five types." Depth is harder to defend because it's interpretive. That doesn't make it less real.

The trap of counting venues instead of experiences

Counting venues is the easiest way to fake rigor. Fifty restaurants, thirty museums, twenty parks—the numbers look impressive until you realize each venue is a single data point, weighted equally, with no sense of what happens inside. A tiny noodle shop where the owner remembers regulars by name scores the same as a tourist-trap chain with a queue around the block. Absurd.

Venue counts measure inventory, not insight. You can inventory a warehouse and still know nothing about what lives there.

— field observation, after a client presented a 400-point dataset with zero field notes

Experiences, by contrast, are layered by nature. They have duration, intensity, and residue—what stays with you after you leave. A one-hour visit to a grand cathedral might produce less residue than a twenty-minute conversation with a bell-ringer who explains why the bells sound different on Tuesdays. That said, experiences are harder to score. They resist quantification. Teams want a number; depth wants a narrative with a number attached.

What usually breaks first is the scoring rubric itself. Someone proposes "experience quality" and the team asks for definitions, metrics, inter-rater reliability—and within a week, they're back to counting venues. I get it. Rankings feel objective; depth feels personal. But personal isn't the same as arbitrary. You can build shared criteria for what counts as a meaningful layer, as long as you accept that it requires discussion, disagreement, and revision.

Start with one neighborhood, not the whole city. Walk the same block three times—morning, midday, night. Write down what changes. That's your first layer stack. Everything else builds from there.

Patterns That Usually Work

Not every benchmark needs a white paper. Some of the most reliable approaches are cheap, repeatable, and surprisingly hard to argue with. Here are four that keep appearing in field notes from teams that actually finish the job.

The walkability test: 20 streets, 3 hours

Pick a residential zone two miles from the tourist core—not the main drag, not the riverfront. Then walk twenty consecutive streets, any direction, and count what actually exists: corner groceries, barbershops, playgrounds, repair stalls, places where someone over sixty is sitting outside. That's your density index. We tried this in Lisbon's Alvalade neighborhood and found eleven cafés that serve food after 9 p.m.; the tourist map promised none. The trick is resisting the urge to check your phone's map for "interesting" routes. Randomness is the point.

What usually breaks first is your patience, not the method. After street number eight, you'll want to head back. Push through—the sixteenth street is where the hidden bakery or the community garden shows up. Depth hides in the boring middle, not the dramatic start. And if you hit three consecutive streets of blank walls and locked gates, that's data too. That's a shallow layer, and knowing that's worth something.

The food-scene density check

Count how many lunch spots within a ten-minute walk serve at least one dish you'd need to ask about to understand. Not exotic—just opaque. A menu with "pork stew, family recipe" is depth. A menu with "asian fusion tacos" is noise. We used this in Bangkok's Ari district and found fourteen places where the owner would happily explain a dish's origin, unprompted. The ratio matters more than the total; five deep options in a hundred restaurants says something different than five in twenty.

That sounds fine until you realize the check has a flaw—it rewards dense, walkable food scenes and punishes spread-out cities where the good stuff is a ten-minute scooter ride away. So adjust: use a bike or a bus for the same radius. The pattern survives the transport change, provided you're honest about the mode shift. The catch is that you'll eat a lot of mediocre meals. That's the cost of sampling honestly.

The nightlife temperature gauge

Go out on a Tuesday, not a Friday. Fridays are performance; Tuesdays are truth. Walk past the same five bars twice—once at 8:30 p.m., once at 11:45 p.m. Count how many venues have different people in them versus the same diehards. High turnover means locals switching spots, which means a living scene. Same faces everywhere means a village, which can be charming, but it's not depth—it's repetition.

One reliable signal is the late-evening staff shift. When bartenders swap out at 11 p.m. and the new crew knows the regulars by name, you're in a place with layers. When the same lone server works the whole night, that's a surface town. Wrong order of checks, though, and you'll misread everything—if you start at the loudest club, you'll think the city is all tourists and neon. Start quiet, end loud, and the gradient tells the real story.

The hidden-corner ratio

'Depth is not how many landmarks you can name; it's how many corners you can find that have no name at all.'

— field note from a long-term traveler, after three years of city-hopping

Take any map app, drop a pin, and screenshot the surrounding blocks. Then go find the dead ends, the alleys that don't connect, the courtyards that look private but aren't. Count them. A city with one hidden corner per ten visible streets has texture—it rewards off-route wandering. A city with zero dead ends is a grid that's been sanitized, and no amount of colorful murals will fix that.

The pitfall here is confirmation bias. If you love a city, you'll find hidden corners everywhere—even in a parking garage. So set the rule before you arrive: you need at least three genuinely surprising spaces in two hours, or the ratio fails. We've seen travelers gush over a single tucked-away plaza and call it depth; that's a sample size of one. The ratio needs volume to mean anything.

What we fixed after early tests was the scoring scale. A hidden corner that leads to a dead end counts half as much as one that opens into a new square or a river view. The first is just a gap; the second is a threshold. That distinction slashed false positives in Vancouver and elevated a few unfamous neighborhoods in Mexico City. Try the same weighting on a short trip—you'll stop seeing "cute spots" and start seeing structure.

Anti-Patterns and Why Teams Revert to Rankings

Rankings are safe. Depth is not. That's the core tension driving most abandonments, and it plays out in predictable ways.

Scoring everything and averaging the noise

The most common failure I have seen is the kitchen-sink scorecard. Teams collect forty metrics—sidewalk width, cafe density, mural count, noise complaints, bike racks per capita—then average them into one tidy number. That tidy number is meaningless. A district with perfect sidewalks and zero murals scores the same as one with crumbling curbs and a wall of street art. You have erased the texture you were trying to measure. Depth is not a single scalar; it's a profile, and profiles have shapes, not totals.

The catch is that averages feel scientific. They produce clean spreadsheets and defensible decisions. But they reward the bland: a place that's okay at everything beats a place that's extraordinary at three things and terrible at two others. And that's precisely the wrong instinct for depth, which thrives on contrast. The fix is to stop aggregating and start weighting—or better, to publish the raw dimensions and let readers form their own composite.

Chasing check-in counts and Instagram tags

Then there is the data proxy trap. When direct observation feels slow, teams reach for whatever numbers are already sitting in an API. Check-in counts, hashtag frequency, review volumes. These are not measures of depth; they're measures of popularity. A tourist-packed plaza with a single chain coffee shop generates thousands of tags. A hidden courtyard with three hole-in-the-wall bakeries and a hand-painted sign generates almost none. Guess which one has more layers.

Odd bit about travel: the dull step fails first.

That sounds fine until someone builds a dashboard on those proxies and calls it depth. It's not depth—it's noise wearing a lab coat. What usually breaks first is the credibility of the whole system. Locals see a top ranking go to a place they would never describe as deep, and the benchmark dies overnight. The alternative is slower: field notes, repeated visits, interviews with people who actually live there. Ugly and expensive, but honest.

Why teams abandon depth after the first pivot

Organizations don't revert to rankings because rankings are better. They revert because rankings are cheap and fast. A pivot arrives—new leadership, a funding squeeze, a quarterly review—and suddenly the depth framework looks like overhead. The maintenance burden is real: scores need rechecking, weights need renegotiating, edge cases need adjudication. A simple rank list doesn't ask anything of anyone.

Depth benchmarking dies not from complexity but from convenience—the ease of a numbered list when context feels like a luxury.

— Observation from a city-data consultant who watched three pilot programs fold

The deeper issue is accountability. Rankings let people off the hook: you can blame the algorithm or the data source. Depth scoring, by contrast, forces judgment calls into the open. Someone has to decide whether a weekly flea market counts more than a permanent bookstore. That's uncomfortable. And when the first dispute hits, teams often discover they never agreed on what "deep" meant in the first place. So they slide back to the ranked list, trading nuance for peace. The path back requires naming that discomfort early and building dispute resolution into the process itself—not pretending the disagreements will vanish.

Maintenance, Drift, and the Long-Term Costs

A depth score is not a one-and-done artifact. It decays, quietly, the moment you stop feeding it fresh observations. The only question is how quickly you notice.

How often to re-measure a neighborhood

Depth benchmarks rot on a schedule, and that schedule is shorter than most teams admit. A city block that felt layered in March can flatten by October—new construction swallows a hidden courtyard, a beloved dive bar becomes a vape shop, the street performer who gave the alley its texture moves on. I have watched teams mark a neighborhood "deep" and then defend that verdict for two years, long after the place had turned generic. Re-measuring every quarter feels excessive until you miss one shift and your entire map quietly lies to users.

The honest cadence depends on how fast the ground changes. Tourist-heavy districts demand monthly checks; stable residential zones can stretch to six months. But here is the pitfall—teams rarely schedule re-measurement at all, because it means pulling people off feature work. That trade-off gets deferred until a user complains, and by then the benchmark is already a parody of what it claims to measure.

Data drift: when your benchmarks go stale

Drift is sneaky. The raw signals you coded into a depth score—photo density, review vocabulary, opening hours variety—all shift independently, and none of them sends a notification. A neighborhood can lose its independent bookshop while the coffee-shop count stays flat, leaving the composite score unchanged but the lived reality hollowed out. What usually breaks first is the weight assignment: something that mattered in 2021 has become noise by 2025, yet your scoring engine still treats it as gospel.

I have seen teams fix this by adding a simple decay function to every underlying metric. Old data points lose influence automatically, so a benchmark that has not been refreshed in a year starts to fade on its own. That approach costs little to build but demands discipline in monitoring—otherwise you just trade a stale benchmark for a vague one. The catch is that decay rates themselves need tuning, and getting them wrong produces either paranoid ratings or complacent ones.

“A depth score is a promise about the present, and the present keeps renegotiating the terms.”

— field note, urban data team lead

The hidden cost of keeping a depth system alive

Nobody budgets for the boring parts. The monthly re-scoring cycles, the edge-case fixes when a new neighborhood gets annexed or a mall renames itself, the occasional full re-platforming when your data vendor changes an API. These costs compound quietly, and they rarely appear in any roadmap review. One team I consulted spent forty hours a quarter just reconciling their benchmark scores against on-the-ground photos—time they never counted in their initial proposal.

The trade-off is brutal but clear: a depth system that's not maintained becomes worse than a simple ranking, because it carries the illusion of truth. Rankings at least admit their crudeness. A stale depth score pretends to know a city's soul and gets more wrong every week. Before committing to this approach, ask who owns the maintenance calendar and what work they will stop doing to keep the benchmarks honest. If the answer is vague, you're building a museum piece, not a tool.

When Not to Use This Approach

Depth scoring is a heavy instrument. Sometimes you need a butter knife. Here's how to tell the difference before you waste three weeks.

When you just need a shortlist, not a thesis

I once spent three weeks building a depth-scoring model for a client who wanted to pick one hotel for a corporate retreat. Twenty-eight metrics. Weighted cultural indices. A geospatial layer for neighborhood vibrancy. The answer was the same hotel they always used—the one with a conference room that fit forty people and a bar that stayed open late. Depth scoring told them nothing they didn't already know, except that they were now three weeks behind on booking.

That's the real test. If someone asks you for a ranking and you hand them a framework, you've failed the assignment. Shortlists work because they're cheap, reversible, and easy to defend. A depth benchmark is a commitment—it implies you're measuring something durable about a place, not just solving a Tuesday problem.

When your audience wants simple answers

Some stakeholders need a number that fits in a slide. They don't want to hear about interaction weights or temporal decay functions. They want "this city is better than that city" and they want it in bold, forty-point type.

Depth scoring produces nuanced outputs: "Osaka scores high on nightlife density but low on third-place accessibility for families." That's useful. It's also unquotable. The moment you feed nuance into a culture that runs on headline metrics, the nuance gets flattened into a rank anyway—except now it's a rank with a fake precision score attached. You haven't improved decision-making; you've just added a coat of paint to the same gut reaction.

What usually breaks first is trust. When someone asks "so, which one wins?" and you answer "it depends on your parameters," they hear "you don't know." That's a fair criticism, actually. If your benchmark can't produce a clear winner for the specific decision at hand, it's not a decision tool—it's a research paper with a waiting list.

When data is too thin to support depth scoring

Depth benchmarks are data hogs. They need reliable streams for foot traffic, event frequency, business churn, public space quality—and they need those streams at a resolution that actually distinguishes between two adjacent neighborhoods. If your data source is one scraper hitting a tourism board's outdated API, you're building a cathedral on sand.

Field note: travel plans crack at handoff.

I have seen teams bolt depth scoring onto datasets that cover three years and call it longitudinal. Three years is not longitudinal. It's a snapshot with a nice haircut. The scoring will drift the moment you feed in a new year's data, and your stakeholders will lose confidence faster than they lost confidence in the old ranking system—which, at least, had the decency to be transparently shallow.

Depth scoring without data coverage is just an expensive way to be wrong with confidence.

— product manager, urban analytics platform

The catch is that thin data makes the scoring look more objective than it's. A single missing variable—say, evening economy hours—can silently shift a city down three slots, and nobody will notice until the results contradict local knowledge. Then you're not just wrong; you're mechanically wrong, which is harder to argue with and harder to fix.

If your dataset can't answer "what changed in this neighborhood last quarter?" without a manual spreadsheet update, depth scoring is premature. Fix the data pipeline first. Reserve the benchmark for places where the signal is dense enough to matter. And if you're ever tempted to force depth scoring onto a shortlist problem again—don't. A two-column table with pros and cons will outperform your beautiful framework every time, because it lets people make a call and move on.

Open Questions and Honest Answers

Let's address the questions that keep circling in planning meetings. No tidy answers here—just trade-offs and what usually shakes out in practice.

How many benchmarks is enough?

Three feels thin. Twelve feels like you're building a dashboard nobody opens. The honest answer: you need enough to catch contradictions, not enough to drown in them. I have watched teams collect ten metrics and then spend weeks arguing about which one is wrong—when the real problem was that two benchmarks measured the same thing from different angles. Start with four: one for movement, one for money, one for culture, one for lived experience. If two of those tell opposite stories, you have found something worth investigating. If they all agree, you have probably built a ranking in disguise.

The catch is that the fourth one—lived experience—is the hardest to define. Foot traffic counts bodies. Spending tracks wallets. Culture is fuzzier. What do you measure when a city's soul is a late-night noodle stand and a mural that got painted over? That's where the field still lacks consensus. Some teams use social media check-ins; others use event density. Neither is right. You're looking for a proxy, not a truth.

Does foot traffic or spending matter more?

Short answer: foot traffic tells you where people want to be; spending tells you what they will actually pay for. They diverge more than you would expect. A weekend market might pull thousands of visitors who buy nothing but coffee. A quiet side street with three high-end boutiques generates more revenue per square meter than the crowded plaza—yet the plaza is what makes the city feel alive. Most teams default to spending because it's measurable. That's a mistake. You lose the texture of a place when you only count receipts.

The practical compromise is to weight them differently per district. Tourist zones lean on spending; residential neighborhoods lean on foot traffic. Wrong order and you end up ranking a mall above a public square—technically accurate, practically useless. I have seen this happen, and the maintenance burden is real: every quarter, someone has to re-check whether the weights still make sense. That friction is why teams revert to rankings. Rankings are lazy. They require no judgment calls.

Can you compare cities across different cultures?

Not cleanly. And pretending otherwise creates more noise than signal. A city where people socialize in public plazas will score higher on foot traffic benchmarks than one where hospitality happens in private homes—but that says nothing about which community is more connected. The fix is not to abandon cross-city comparisons. The fix is to compare patterns, not absolute numbers. Look at how a metric changes over time within each city, then compare the shapes of those changes. A rising curve in Tokyo and a rising curve in Mexico City might rhyme even if their baselines look nothing alike.

Depth scoring works when you ask "how does this place change?"—not "which place is better?" The former teaches you something; the latter just feeds ego.

— framework note, distilled from field observations

That said, some cultural dimensions resist quantification altogether. Religious festivals, seasonal rhythms, the way people queue—these shape a city's depth, but they don't fit into benchmarks. Flag them as qualitative context. It's messy. It's incomplete. But it beats the false confidence of a single number. The next experiment worth running: pick one district, track your four benchmarks for two weeks, and write down what the numbers miss. Compare that list to what the numbers captured. That gap is where the real work lives.

What to Try Next: A Few Experiments

Enough theory. Here are three concrete things you can run this month, with minimal setup and a decent chance of surprising yourself.

Run a one-day pilot on a single district

Pick one neighborhood you know well—maybe the one you always tell visitors about. Map out three depth signals you can actually observe by foot: how many independent shops survive past 6pm, how many elders sit in public spaces, how many street names reference local history. Walk it for an hour. Count what you see. The first pass feels awkward; you will miss half the signals. That's the point.

Next morning, do the same walk with a different lens. Look for what is missing instead of what exists. No benches near bus stops. No public toilets. No shadows in the afternoon heat. That absence list often tells you more about a place than any presence count. Most teams stop after the first walk—they collect data, then rush to rank. Resist that.

Compare your depth score with a ranking list

Take your raw notes and convert them into a crude score—say, 1 to 5 per signal. Then pull up a standard city ranking from any travel site. Put them side by side. The mismatch will sting. A top-ten city might score 2.4 on depth, while a "boring" industrial town hits 4.8. That gap is not an error; it's the finding. Write down three reasons the ranking and your score disagree. Those reasons become your thesis.

One caveat: don't over-engineer the scoring on day one. Fractions and weights and normalization curves will eat your afternoon and produce false confidence. A simple sum works. The goal is to see where your intuition about a place clashes with the aggregated opinion of millions of tourists. That friction is the raw material.

Interview a local and track what they mention

Sit with a shopkeeper, a retired teacher, or a bus driver. Ask one question: "What changed here in the last ten years?" Don't prompt them with categories. Just listen. Track how many times they mention physical infrastructure versus social routines versus economic pressure. Notice what they do not mention—often the most telling data point. I have done this in three cities now, and the pattern repeats: locals describe depth through daily rhythms, not landmarks.

The tricky bit is resisting the urge to steer the conversation toward what you already measured. Let them wander. Their detours are the real benchmark. Compare their mental map to your walking notes—the overlap is your depth signal, the divergence is your blind spot.

Ask ten people what makes their city livable. You will get ten different answers, all circling the same hidden thing.

— field note, after a week of interviews in Rotterdam

Run these three experiments within a month. Keep the notes messy. Compare them, then throw away the scoring system and rebuild it from what the interviews revealed. The ranking comparison will feel dated by week two. That's fine—depth benchmarking is a practice, not a product. The next iteration will always beat the current one.

Share this article:

Comments (0)

No comments yet. Be the first to comment!