doops in the wind · a working sketchbook

Dina Pisareva

04 · The Claude Playbox · a method, explained

Word Embeddings as a Ruler

How three sociologists measured what a word means by subtracting one word from another, how their ruler differs from the embeddings a chatbot vendor sells, what it has already shown about disease and stigma, and what it could do with patient records. Written by Claude, Anthropic’s model, at Dina’s request for a co-author in public health, with twenty doodles.

01Every word gets an addressa guessing game, played a few billion times

A word embedding gives every word in a pile of text an address made of a few hundred numbers, and the address is learned by playing one guessing game over and over.

The best-known tool is word2vec, published by Tomas Mikolov and colleagues at Google in 2013. It reads text through a small sliding window, five to ten words wide. At each stop the model uses the middle word to guess its neighbors, or the neighbors to guess the middle word. Every word starts with 300 random numbers, which together are its vector. When a guess is wrong, which at first is always, the model nudges the numbers of the word and its real neighbors a little closer together, and nudges a few random words a little further away. Nobody tells the model what any word means. After a few billion windows the numbers settle, and the settled numbers are the map.

Plate 1The guessing game
THE GUESSING GAME, PLAYED A FEW BILLION TIMEStherichbankercountedhismoney1. a window of fiveslides over the textbanker? thenrich, counted...2. guess the neighborsfrom the wordbankerrichlacrossecloser to its neighbors,away from a random word3. nudge, repeat
How word2vec learns. A window slides over the text, the model guesses the neighbors of the middle word from its numbers, and after each guess it nudges the word toward its real neighbors and away from a few random ones. Billions of nudges later, the numbers are the map.after Mikolov et al., 2013

Two words do not need to meet to become neighbors on the map. If “diabetes” and “hypertension” both keep appearing near “chronic,” “managed” and “medication,” they receive the same nudges and land close together, even if no sentence mentions both. The linguist J. R. Firth stated the principle in 1957: you shall know a word by the company it keeps. The map never reads a definition. It only counts company.

Plate 2Known by the company
EVERY SHARED CONTEXT IS A NUDGEPOORRICHbankeraffluentmoneyedrichnannyneedydestitutepoorThe map never reads a definition. It only counts neighbors.
The example Kozlowski, Taddy and Evans use. Every shared window with “affluent,” “moneyed” or “rich” nudges “banker” toward the rich end, and every one with “needy,” “destitute” or “poor” nudges “nanny” the other way.after Kozlowski, Taddy and Evans, 2019, p. 906

The map measures closeness with cosine similarity, the cosine of the angle between two words’ arrows. A cosine of 1 means the arrows point the same way, 0 means they meet at a right angle and have nothing to do with each other, and −1 means they point in opposite directions. One fact about high dimensions matters for everything that follows. Pick two directions at random in a space with 300 axes and they will almost always meet at very nearly a right angle. A cosine of 0.3, which would be a shrug on a flat page, is therefore a strong statement in 300 dimensions. Right angles are the default, and anything far from a right angle was put there by the text.

02Rich minus poorthe subtraction that makes a ruler

In 2019 three sociologists turned the map into a measuring instrument by subtracting one word from another, and that subtraction is the whole method.

Austin Kozlowski, Matt Taddy and James Evans published The Geometry of Culture in the American Sociological Review in 2019, volume 84, pages 905 to 949. Their starting point was a party trick that word2vec was already famous for. Take the vector for “king,” subtract “man,” add “woman,” and the nearest word to the result is “queen.” The trick works because the arrow from “man” to “woman” is the same arrow as from “king” to “queen.” The direction of a difference between two words carries a meaning of its own.

Plate 3Word arithmetic
WORD ARITHMETICmanwomankingqueenking − man + woman ≈ queenPOORRICHhockeylacrosse+ affluence − povertyfancy.
Left, the famous parallelogram: king minus man plus woman lands on queen. Right, the sociologists’ version: take “hockey,” add the direction from poverty to affluence, and land next to “lacrosse.”after Mikolov et al., 2013, and Kozlowski, Taddy and Evans, 2019

The authors built a cultural dimension from that observation. Take a pair of antonyms that name the distinction you care about, such as rich and poor, and subtract one vector from the other. The result is an arrow pointing from poor toward rich. A single pair is noisy, since “rich” also means a chocolate cake, so they collected 42 such pairs from five thesauri, affluent and impoverished, opulent and indigent, expensive and cheap, and averaged the 42 arrows into one. That average is the affluence dimension. Gender came from pairs such as woman and man, she and he, and a dozen other distinctions were built the same way: education, status, cultivation, morality, employment, race.

Plate 4The class ruler
THE CLASS RULER42 pairs, one average, one directionrich − pooropulent − indigentritzy − ramshackleswanky − basicflush − skint...37 moreAFFLUENCEPOORRICHboxingcampingbaseballbasketballvolleyballtennisgolfprojection = cosine between the word and the yellow arrow
Dozens of small arrows, rich minus poor, opulent minus indigent and so on, average into one fat arrow labeled affluence. A word’s projection is the cosine between its own arrow and that one.after Kozlowski, Taddy and Evans, 2019, Figure 1

Any word can then be laid against the dimension. Its projection is the cosine between the word’s vector and the dimension’s arrow, a number between −1 and 1, in practice between about −0.3 and 0.3. On the Google Books affluence dimension, golf and tennis project rich, boxing and camping project poor, and baseball sits near zero. You can compute the same projection for a food, a car, a job, a first name or a disease. The dimension is a ruler, and the projection is a reading on it.

Plate 5One pair is noisy
ONE PAIR IS NOISYrich − poor, alone: r ≈ .2 with people42 pairs averaged: r ≈ .5, and 10 did about as well
A dimension built from one antonym pair correlates about .2 with human ratings of class. Averaging many pairs steadies the arrow, and the gains flatten after about ten pairs.after Kozlowski, Taddy and Evans, 2019, Appendix C

One careful sentence about what a projection means. A cosine of 0.1 on affluence does not mean a word is 10 percent rich. It means that on this map the word leans slightly toward the rich end of a direction that most words meet at a right angle. Readings are comparable within one model and one dimension. “Golf projects richer than boxing in these books” is a fair sentence. “Golf is 0.12 rich” is not a sentence at all. Across two different models the raw numbers are not comparable, because each model has its own spread, so the paper compares orders and gaps rather than values, and so should anyone who borrows the ruler.

03Checking the ruler against peoplea steak, rated by 398 strangers

Before trusting the ruler on a century of dead authors, the authors checked it on the living, by paying people to rate a steak.

They posted a survey on Amazon Mechanical Turk, a site where people do small paid tasks, in 2016 and again in 2017. In all, 398 people from the United States finished it, for $1.75 each. Each rated 59 words from 0 to 100 on three scales: very working class to very upper class, very feminine to very masculine, and white to African American. The words came from seven domains, among them occupations, foods, clothing, sports and first names. The authors weighted the answers to match the national population by sex, education and race.

Plate 6Rate the steak
RATE THE STEAKHow would you rate a steak?very working classvery upper class398 Mechanical Turk workers, 59 words, 0 to 100, $1.75 eachHOW WELL THE MAP AGREED WITH THEMclassr = .53genderr = .76racer = .27the Google Books model, against the survey averages
The survey question and the result. For class, the Google Books model’s projections correlated .53 with the survey averages, for gender .76, for race .27. Two models trained on news and web text did better on class and gender and worse on race.after Kozlowski, Taddy and Evans, 2019, Table 1

The projections and the survey averages agreed to a degree that depends on the dimension. For class the correlation was .53 on the Google Books model and .58 and .57 on two models trained on news and the open web. For gender it was .76, .88 and .90. For race it was .27, .75 and .44. A second test asked a humbler question: of any two words, does the model put them in the same order as the people? For class the Books model got 85 percent of pairs right, for gender 75 percent, for race 69 percent. The lesson is that every dimension has to earn its own validation. The gender ruler is sharp, the class ruler is usable, and the race ruler built from books deserves no trust on its own.

Others have checked the ruler against records rather than surveys. In 2018 Nikhil Garg and colleagues at Stanford measured, in a Google News embedding, how feminine each occupation word projected, and set that against the share of women in each occupation in the census. The two agreed with an r-squared of .46, and the fitted line ran through the origin, so an occupation split fifty-fifty showed no bias either way. A year earlier Aylin Caliskan, Joanna Bryson and Arvind Narayanan had found a correlation near .90 between the gender association of 50 occupation words and the Bureau of Labor Statistics share of women in them. Language carries the statistics of the society that wrote it, down to the payroll.

Plate 7The map against the census
THE MAP AGAINST THE CENSUSshare of women in the occupation, censushow feminine theword projectsnurselibrarianteachercashieraccountantphotographerchemistengineercarpenterr² = .46, line through zeroillustrative points. the fit statistic is the paper’s
Garg and colleagues set each occupation’s embedding bias against the census share of women in it. The points here are illustrative. The fit statistic is the paper’s.after Garg, Schiebinger, Jurafsky and Zou, PNAS, 2018, Figure 1A

04Angles, error bars and timethe part the other methods cannot do

The ruler gives three things that word counts cannot: the angle between two distinctions, an error bar on every reading, and the same measurement repeated across decades.

Two dimensions are themselves arrows, so you can measure the angle between them. The authors found that cultivation and employment meet at 90 degrees, status sits at 85.5 degrees from cultivation and at 79.3 degrees from employment. No flat page can hold those three angles, since 85.5 plus 79.3 is more than 90, and that is the paper’s plain demonstration that class in American books is not one thing. For a clinic the same measurement would ask whether the blame distinction in notes leans toward the race distinction, the insurance distinction or neither.

Plate 8No room on a flat page
NO ROOM ON A FLAT PAGEemploymentcultivation90°status85.5° from cultivation79.3° from employment85.5 + 79.3is more than 90.I need to leave.the third arrow has to come up out of the paper
Cultivation and employment sit at a right angle. Status must be 85.5 degrees from one and 79.3 from the other, which no flat drawing allows. The third arrow has to come up out of the paper.after Kozlowski, Taddy and Evans, 2019, Table 3

The error bars come from a trick that any corpus allows. The authors cut the Google Books text into 20 pieces, trained a model on each, and read the projection off all 20. The band that holds 90 percent of the 20 readings is the error bar. A word that appears fewer than about 25 times wobbles so much that the band swallows the reading, and the paper drops such words from the map. A rare word is not a small signal. It is no signal.

Time comes from training one model per decade and laying the same ruler on each. Across the twentieth century the cosine between the education dimension and the affluence dimension climbed from about .15 in the 1900s to about .42 in the 1990s, while the other distinctions held their angles. In the books, schooling became what money means. The same design would serve a hospital whose notes span 20 years, with one model per year or per five-year block and the same antonym pairs each time.

Plate 9Education climbs
EDUCATION CLIMBScloseness to affluence, cosine, Google Books by decade1900s1920s1940s1960s1980s1990s0.00.20.4educationcultivationstatusmoralityemploymentgendershapes read off Figure 5, not its exact values
The cosine between education and affluence rises decade by decade while cultivation, status, morality and employment hold roughly flat and gender stays near zero. Shapes read off the paper’s Figure 5, not its exact values.after Kozlowski, Taddy and Evans, 2019, Figure 5

05Five habits before you believe a numberthe ruler is yours, the reading is only as good as the ruler

Everything the method can get wrong follows from one fact: it measures how a community of writers used words, and nothing else.

Plate 10Five habits
FIVE HABITS BEFORE YOU BELIEVE A NUMBERwhose corpus?books by a literary public, not everybodywords, not people"nurse" leans feminine in print, no nurse was askedcheck the bandrare words wobble, under 25 uses is off the mappairs firstwrite the antonyms down before you looka fluke is a fluke"patient" near employee until a second corpus agreesthe ruler is yours.the reading is onlyas good as the ruler.
Ask whose corpus it is. Remember that words were measured, not people. Check the error band. Write the antonym pairs down before you look at any result. Treat a single surprising reading as a fluke until a second corpus agrees.after Kozlowski, Taddy and Evans, 2019, and Arseniev-Koehler, 2024

First, whose corpus. Google Books is written by people who publish books, and a hospital’s notes are written by the people who write notes. The map is of the writers’ culture. Second, words, not people. If “nurse” projects feminine, that is a fact about print, and no nurse was asked. Third, the band. Report the 90 percent band with every reading, and do not read a word that falls below the frequency floor. Fourth, pairs first. The antonym pairs define the ruler, so write them down, and the reason for each, before looking at any projection, since a ruler chosen after the result will always confirm it. Fifth, flukes. A single reading that surprises you is a lead to check in a second corpus, not a finding. Alina Arseniev-Koehler’s 2024 review in Sociological Methods & Research walks through these choices one by one and is the best single thing to read after the paper itself.

06Three kinds of embeddingsword2vec is not the thing you send to OpenAI

The word “embedding” now names three different objects, and only the first one is a ruler.

Word2vec gives one fixed address per word type. “Bank” has one vector, parked halfway between the river and the money, and it is the same vector every time. That is a weakness for reading a sentence and a strength for measuring a culture, since the vector is a summary of every context the word ever had in that corpus.

A transformer, the kind of model inside a chatbot, gives a new vector for each use of a word, at each of many layers. The “bank” in “we sat on the river bank” and the “bank” in “the bank refused the loan” receive different vectors, because every layer rewrites the word in the light of its neighbors. These contextual embeddings are what make the model good at reading. They are harder to use as a ruler, because there is no single address for a word to project, though two recent papers from Kozlowski’s group show it can be done by averaging a word’s hidden states over several prompts.

Plate 11One address, or a new one every time
ONE ADDRESS, OR A NEW ONE EVERY TIMEword2vecbankrivermoneyone pin, fixed,halfway between botha transformerwe sat on the river bankthe bank refused the loana new pin for each use,at each of many layersan embeddings APIa whole paragraphabout mothers-in-lawtrained on ?1,536 numbers, tunedto find similar passages,not to measure a word
Word2vec parks one pin per word. A transformer drops a new pin for each use at each layer. An embeddings API returns one vector for a whole passage, trained to find similar passages, from a corpus the vendor does not disclose.after Mikolov et al., 2013, Vaswani et al., 2017, and the OpenAI embeddings documentation

When people “send text to OpenAI for embeddings,” they get a third thing. The embeddings endpoint takes a whole passage, a sentence or a page, and returns one vector of 1,536 or 3,072 numbers for the passage, not for any word in it. Their makers train them with a contrastive objective: pull together texts that belong together, a question and its answer, a passage and its paraphrase, and push unrelated texts apart. The geometry is tuned to answer one question well, which is “which of these passages are about the same thing,” the question search engines ask. The vendor does not disclose the training text, so nobody can date it or cut it into 20 pieces, and the vendor can replace the model at any time, after which the old numbers no longer compare with the new ones. Such vectors are excellent for finding the 50 notes most like this one. They cannot tell you what a word meant.

Plate 12Where a chatbot keeps its words
WHERE A CHATBOT KEEPS ITS WORDSunembedding: pick the next wordlayer 40 of 80: hidden states...layer 8 of 28: hidden statesembedding matrix: the front desktokenpianoout: “beautiful”2025 paperreads the desk2026 paperreads the floors
A token enters at the front desk, the embedding matrix, where it has a fixed address like a word2vec vector. It then rises through layers of hidden states, rewritten at each floor, and leaves at the top as the next word. Kozlowski, Dai and Boutyline read the desk in 2025. Kozlowski and Boutyline read the floors in 2026.after Kozlowski, Dai and Boutyline, 2025, and Kozlowski and Boutyline, 2026

The two recent papers matter for one reason here. They show that the subtraction travels. Kozlowski, Callin Dai and Andrei Boutyline took the fixed word table at the front of Google’s Gemma models, built 28 antonym dimensions, projected 301 words, and found correlations with human ratings between .3 and .7. Kozlowski and Boutyline then did the same inside the hidden layers of Meta’s Llama models, for 360 words on 32 dimensions, checked against ratings from 1,750 people, and found the best dimensions above .8. They also found that the angles between the dimensions inside the model predict the correlations between the same scales in the human survey, with r of .87. The model’s geometry and the people’s judgments have the same shape. The interpretability people who steer chatbots by adding a direction to a hidden state are doing rich minus poor with a different corpus.

Plate 13The same subtraction
THE SAME SUBTRACTION2019, a libraryrichpoorrich − pooraveraged over 42 pairs2025, a chatbotvillain linesneutral linesvillain − neutraladded while the model writes
In 2019 the affluence arrow was rich minus poor, averaged over 42 pairs and read off a library. In 2025 a persona arrow is villain minus neutral, read off a chatbot’s hidden states and added back while it writes.after Kozlowski, Taddy and Evans, 2019, and Chen et al., 2025

07Health, measured this way alreadydisease, stigma and the intensive care unit

The ruler has been laid on health five times already, in newspapers and in clinical notes, and each study shows one thing it can do.

Alina Arseniev-Koehler and Jacob Foster trained word2vec on the New York Times and built four rulers: gender, morality, health and class. They placed words about body weight on each. Fat projected immoral, unhealthy, poor and feminine, and the four readings moved separately. The newspaper carried four meanings of fat, and the method could say which was strongest.

Rachel Kahn Best and Arseniev-Koehler did the same for 106 diseases in 4.7 million news articles from 1980 to 2018. They built rulers for immorality, bad character and disgust. Preventable and behavioral conditions, such as addiction and obesity, sat at the immoral end. Infectious diseases sat at the disgust end. Over four decades stigma fell for chronic physical illness and for nothing else. One chart holds all 106 conditions, and that is the design to copy.

Plate 14106 diseases on two rulers
106 DISEASES ON TWO RULERSnews articles, 1980 to 2018LESS DISGUSTMORE DISGUSTLESS IMMORALMORE IMMORALdiabetescancerdepressionHIVflucancerarthritisHIVaddictionobesityinfectious diseases lean right on the top rulerbehavioral and preventable conditions lean right on the bottom onepositions are illustrative, the two patterns are the paper’s
Best and Arseniev-Koehler placed 106 health conditions on rulers of disgust and immorality, built from news text from 1980 to 2018. Infectious diseases leaned toward disgust, behavioral and preventable conditions toward immorality. Positions here are illustrative.after Best and Arseniev-Koehler, ASR, 2023

Julien Cobert and colleagues trained word2vec on intensive care notes from two hospitals, San Francisco from 2012 to 2022 and Boston from 2001 to 2012. They measured how close the words for Black and White patients sat to words for violence and noncompliance. In Boston the Black words sat closer to violence. In San Francisco the White words did. The notes of two hospitals carry two different cultures, and a model trained on either carries it too. Haoran Zhang and colleagues showed the same for BERT trained on the MIMIC-III notes: a model built on the notes performed differently by gender, language, ethnicity and insurance.

Two studies without embeddings show what the words in a note do. Anna Goddu and colleagues gave 413 physicians in training a chart note about a 28-year-old man with sickle cell disease. Half read a note that called him difficult, half read a neutral note about the same facts. The first half liked him less and treated his pain less aggressively. Michael Sun and colleagues searched 40,113 notes at one hospital for words such as refused, not adherent and combative. Black patients were two and a half times as likely to carry one. A note is a treatment the next clinician receives, and it is not given out evenly.

Plate 15The words change, the mood does not
THE WORDS CHANGE, THE MOOD DOES NOT1850slonelyignorant1900ssavagepoor1950suntidylazyhalf the top traits turn over each decade, the valence staysillustrative trait words. the pattern is Charlesworth et al., 2022
Across 200 years of Google Books, the traits attached to a social group turn over by half each decade while the overall feeling stays put. A time series needs both: which words, and how warm.after Charlesworth et al., PNAS, 2022

08When the corpus is smallyou do not need a hundred million words

Kozlowski had a century of books. Most health projects have a few thousand documents, and three tools make the ruler work at that size.

Facebook’s fastText project published static word vectors for 157 languages, Russian and Kazakh among them, trained on Wikipedia and a crawl of the web, 300 numbers per word, built from pieces of words so that a noun with twelve suffixes shares its evidence across all twelve forms. For biomedical English there is BioWordVec, published in Scientific Data in 2019 and trained on PubMed abstracts and the MIMIC-III notes. Pretrained vectors carry the full Kozlowski recipe, antonym pairs and all, with no training at all. What they cannot do is tell you how your community used a word, because the map is someone else’s.

Plate 16À la carte
À LA CARTEMENU1. take pretrained vectors (GloVe, 300 numbers a word)2. find every sentence where your word appears3. average the vectors of the words around it4. multiply by a learned matrix A to fix the average5. compare the result between groups, with a permutation testserves a word with twenty mentionssix matrix multiplies, a few milliseconds
A corpus-specific vector from a few mentions. Take pretrained vectors, average the ones around each mention of your word, correct the average with a learned matrix, and compare the result between groups with a permutation test.after Rodriguez, Spirling and Stewart, APSR, 2023

Pedro Rodriguez, Arthur Spirling and Brandon Stewart solved that in 2023 with à la carte embeddings. For every mention of your word in your corpus, average the pretrained vectors of the words around it, then multiply the average by a matrix learned once from a large corpus to undo the blurring. The result is a vector for your word in your text, and it works with fewer than twenty mentions. Their R package, conText, compares the vector between groups, such as notes from two departments or two years, with a permutation test for the difference.

Dustin Stoltz and Marshall Taylor approached from the other end with concept mover’s distance in 2019. Instead of asking where a word sits, it asks how far the words of one document would have to travel across the map to reach a concept, such as death or blame. A short trip means the document engages the concept. It works on a tweet or a novel, with pretrained vectors, and it gives a score per document, which is the natural unit when the documents are notes.

Plate 17The concept mover
THE CONCEPT MOVERa documentgrave tomb mourn sleepthe conceptdeathhow far must the words travel?a short trip means the text engages the conceptworks on a tweet or a novel, with pretrained vectors
How far must the words of one document travel across the map to reach the concept? A short trip means the text engages it. The method scores single documents and runs on pretrained vectors.after Stoltz and Taylor, 2019

09A note from Claude on patient recordspain is the best case for the ruler

Dina asked me what the ruler could do with health data. This section is my answer, and it is signed.

In a hospital the ruler measures the clinicians. Notes are written by clinicians, so a map built from notes is a map of how clinicians write about patients. That is the opportunity. Goddu showed that the words in a note change what the next reader does, and Sun counted the words. The ruler adds the angle. Build a blame ruler from pairs such as compliant and noncompliant, cooperative and combative. Build an insurance ruler and a race ruler from the words the notes use. Measure the angle between blame and each. If blame leans toward the uninsured pole and stands at a right angle to race, the notes have said where their moral vocabulary lives. Put 100 diagnoses on the blame ruler and the chart shows which conditions carry it.

Plate 18The unit is the corpus
THE UNIT IS THE CORPUS40,000 notesby many handsone mapof how they wroteBLAMEDEXCUSEDa number per wordnever per patientnot a reading on anyone
Forty thousand notes become one map. The map yields one reading per word on a ruler. No reading belongs to any patient, and the method must never be used to produce one.Claude, 2026

The number belongs to a word, never to a patient. A projection says how a community of writers used a word. It says nothing about a person in the record, and it must not be turned into a score for one. That sentence goes into the ethics application, and it is also the method’s best defense. Nothing in the output identifies, ranks or predicts anyone.

Pain is the best case for the ruler, because pain exists in medicine only as words. A blood sugar is a number from a machine. Pain is a number the patient picks from 0 to 10 and the words the patient and the clinician use around it. Those words are the data, and the ruler is a method for words. Three measurements follow.

First, a pain-severity ruler. The McGill Pain Questionnaire, written by Ronald Melzack in 1975, is a list of 78 pain words, such as throbbing, stabbing, gnawing and unbearable, sorted into 20 groups and ranked by intensity from human ratings. Those ranks are a ready-made human panel, like Kozlowski’s 398 steak raters. Build a severity ruler from pairs such as mild and severe, bearable and unbearable, and check the projections of the 78 words against Melzack’s ranks. A ruler that passes can score any pain vocabulary at all, including words the questionnaire never listed and words in Russian or Kazakh, where no ranked pain word list exists.

Second, a credibility ruler. Kelly Hoffman and colleagues showed in 2016 that half of a sample of white medical students and residents held false beliefs about Black patients’ bodies and rated their pain lower. In notes that bias has a vocabulary. Patients report, state and endorse pain, or they claim, insist and complain of it. Build the ruler from believed and doubted, reports and claims, and project the words that surround pain in notes by patient group, by department and by year. The 2016 opioid prescribing guideline is a before and after. If pain words moved toward the doubted end after it, the notes show the guideline’s effect on language, with an error band.

Third, the translation gap. A patient writes that her chest is being crushed. The note says chest pressure. Project patient words and chart words for the same complaint on the same severity ruler, and the distance between them measures what the chart lost. Portal messages, complaints and forum posts give the patient side. The same test fits chronic, stable and routine, which clinicians use to reassure and patients hear otherwise.

Plate 19The pain ruler
THE PAIN RULERMILDUNBEARABLEnaggingachingstabbingunbearable78 words,rankedthrobbinggnawingsearingMcGill, 1975DOUBTEDBELIEVEDclaimscomplains ofreportsendorsesthe words around a pain report, on two rulers. positions are illustrative
Two rulers for the words around a pain report. Severity, checked against the 78 ranked words of the McGill Pain Questionnaire. Credibility, from claims and complains of to reports and endorses. Positions are illustrative.Claude, 2026, after Melzack, 1975
Plate 20One word, two maps
ONE WORD, TWO MAPSclinicians’ notesa chartMILDSERIOUSchronicpatients’ messagesMILDSERIOUSchronicchronic meansforever.the same word, projected on the same ruler, built from two corpora
The same word, “chronic,” projected on a mild to serious ruler built from clinicians’ notes and again from patients’ messages. The gap between the two readings is the measurement.Claude, 2026

Notes need cleaning before training, and the cleaning is most of the work. Templates repeat ten thousand times and pull every word in them together, so strip them. Negation puts chest pain next to denies, so mark negated spans, or the model learns that pain is a thing patients deny. Abbreviations need a pass, or SOB lands somewhere strange. Check the nearest neighbors of twenty common symptoms by hand before believing anything.

Validate with two panels. Have clinicians and patients rate the same words on the same scales. The disagreement between the two panels is a finding on its own. Sun’s hand-coded word list is the other test. A stigma ruler that predicts the hand codes has earned trust.

Vectors are patient data. John Morris and colleagues recovered full names from clinical note embeddings in 2023. Train inside the institution, publish projections and bands and never the vectors of rare words, and send nothing to a vendor’s endpoint. The frequency floor the method needs anyway, dropping words seen fewer than 25 times, is also a privacy floor.

Write the pairs down first. The antonym pairs are the hypothesis. A team that registers its pain pairs, blame pairs and social pairs before seeing a projection has a study. A team that picks pairs while looking has a mirror. The first deliverable is one page: the rulers, the pairs behind each, the corpora, and what a null result looks like.

Claude
Anthropic’s model, writing with Dina Pisareva, 7 October 2026

10Reading listin the order the page uses them

  1. Kozlowski, Taddy and Evans, The Geometry of Culture: Analyzing the Meanings of Class through Word Embeddings, American Sociological Review 84(5), 2019
  2. Kozlowski and Boutyline, The Semantic Structure of Feature Space in Large Language Models, arXiv 2604.27169, 2026
  3. Kozlowski, Dai and Boutyline, Semantic Structure in Large Language Model Embeddings, arXiv 2508.10003, 2025
  4. Arseniev-Koehler, Theoretical Foundations and Limits of Word Embeddings: What Types of Meaning Can They Capture?, Sociological Methods & Research 53(4), 2024
  5. Arseniev-Koehler and Foster, Machine Learning as a Model for Cultural Learning: Teaching an Algorithm What It Means to Be Fat, Sociological Methods & Research 51(4), 2022
  6. Best and Arseniev-Koehler, The Stigma of Diseases: Unequal Burden, Uneven Decline, American Sociological Review 88(5), 2023
  7. Cobert et al., Measuring Implicit Bias in ICU Notes Using Word-Embedding Neural Network Models, Chest 165(6), 2024
  8. Zhang, Lu, Abdalla, McDermott and Ghassemi, Hurtful Words: Quantifying Biases in Clinical Contextual Word Embeddings, ACM CHIL, 2020
  9. Goddu et al., Do Words Matter? Stigmatizing Language and the Transmission of Bias in the Medical Record, Journal of General Internal Medicine 33(5), 2018
  10. Sun, Oliwa, Peek and Tung, Negative Patient Descriptors: Documenting Racial Bias in the Electronic Health Record, Health Affairs 41(2), 2022
  11. Garg, Schiebinger, Jurafsky and Zou, Word Embeddings Quantify 100 Years of Gender and Ethnic Stereotypes, PNAS 115(16), 2018
  12. Caliskan, Bryson and Narayanan, Semantics Derived Automatically from Language Corpora Contain Human-like Biases, Science 356, 2017
  13. Charlesworth, Caliskan and Banaji, Historical Representations of Social Groups across 200 Years of Word Embeddings from Google Books, PNAS 119(28), 2022
  14. Rodriguez, Spirling and Stewart, Embedding Regression: Models for Context-Specific Description and Inference, American Political Science Review 117(4), 2023
  15. Stoltz and Taylor, Concept Mover’s Distance, Journal of Computational Social Science 2, 2019
  16. Zhang, Chen, Yang, Lin and Lu, BioWordVec, Improving Biomedical Word Embeddings with Subword Information and MeSH, Scientific Data 6, 2019
  17. Morris, Kuleshov, Shmatikov and Rush, Text Embeddings Reveal (Almost) As Much As Text, EMNLP, 2023
  18. Melzack, The McGill Pain Questionnaire: Major Properties and Scoring Methods, Pain 1(3), 1975
  19. Hoffman, Trawalter, Axt and Oliver, Racial Bias in Pain Assessment and Treatment Recommendations, PNAS 113(16), 2016
  20. Mikolov, Chen, Corrado and Dean, Efficient Estimation of Word Representations in Vector Space, 2013
  21. Grave, Bojanowski, Gupta, Joulin and Mikolov, Learning Word Vectors for 157 Languages, LREC, 2018