doops in the wind · a working sketchbook
04 · The Claude Playbox · a method, explained
Word Embeddings as a Ruler
How three sociologists measured what a word means by subtracting one word from another, how their ruler differs from the embeddings a chatbot vendor sells, what it has already shown about disease and stigma, and what it could do with patient records. Written by Claude, Anthropic’s model, at Dina’s request for a co-author in public health, with twenty doodles.
01Every word gets an addressa guessing game, played a few billion times
A word embedding gives every word in a pile of text an address made of a few hundred numbers, and the address is learned by playing one guessing game over and over.
The best-known tool is word2vec, published by Tomas Mikolov and colleagues at Google in 2013. It reads text through a small sliding window, five to ten words wide. At each stop the model uses the middle word to guess its neighbors, or the neighbors to guess the middle word. Every word starts with 300 random numbers, which together are its vector. When a guess is wrong, which at first is always, the model nudges the numbers of the word and its real neighbors a little closer together, and nudges a few random words a little further away. Nobody tells the model what any word means. After a few billion windows the numbers settle, and the settled numbers are the map.
Two words do not need to meet to become neighbors on the map. If “diabetes” and “hypertension” both keep appearing near “chronic,” “managed” and “medication,” they receive the same nudges and land close together, even if no sentence mentions both. The linguist J. R. Firth stated the principle in 1957: you shall know a word by the company it keeps. The map never reads a definition. It only counts company.
The map measures closeness with cosine similarity, the cosine of the angle between two words’ arrows. A cosine of 1 means the arrows point the same way, 0 means they meet at a right angle and have nothing to do with each other, and −1 means they point in opposite directions. One fact about high dimensions matters for everything that follows. Pick two directions at random in a space with 300 axes and they will almost always meet at very nearly a right angle. A cosine of 0.3, which would be a shrug on a flat page, is therefore a strong statement in 300 dimensions. Right angles are the default, and anything far from a right angle was put there by the text.
02Rich minus poorthe subtraction that makes a ruler
In 2019 three sociologists turned the map into a measuring instrument by subtracting one word from another, and that subtraction is the whole method.
Austin Kozlowski, Matt Taddy and James Evans published The Geometry of Culture in the American Sociological Review in 2019, volume 84, pages 905 to 949. Their starting point was a party trick that word2vec was already famous for. Take the vector for “king,” subtract “man,” add “woman,” and the nearest word to the result is “queen.” The trick works because the arrow from “man” to “woman” is the same arrow as from “king” to “queen.” The direction of a difference between two words carries a meaning of its own.
The authors built a cultural dimension from that observation. Take a pair of antonyms that name the distinction you care about, such as rich and poor, and subtract one vector from the other. The result is an arrow pointing from poor toward rich. A single pair is noisy, since “rich” also means a chocolate cake, so they collected 42 such pairs from five thesauri, affluent and impoverished, opulent and indigent, expensive and cheap, and averaged the 42 arrows into one. That average is the affluence dimension. Gender came from pairs such as woman and man, she and he, and a dozen other distinctions were built the same way: education, status, cultivation, morality, employment, race.
Any word can then be laid against the dimension. Its projection is the cosine between the word’s vector and the dimension’s arrow, a number between −1 and 1, in practice between about −0.3 and 0.3. On the Google Books affluence dimension, golf and tennis project rich, boxing and camping project poor, and baseball sits near zero. You can compute the same projection for a food, a car, a job, a first name or a disease. The dimension is a ruler, and the projection is a reading on it.
One careful sentence about what a projection means. A cosine of 0.1 on affluence does not mean a word is 10 percent rich. It means that on this map the word leans slightly toward the rich end of a direction that most words meet at a right angle. Readings are comparable within one model and one dimension. “Golf projects richer than boxing in these books” is a fair sentence. “Golf is 0.12 rich” is not a sentence at all. Across two different models the raw numbers are not comparable, because each model has its own spread, so the paper compares orders and gaps rather than values, and so should anyone who borrows the ruler.
03Checking the ruler against peoplea steak, rated by 398 strangers
Before trusting the ruler on a century of dead authors, the authors checked it on the living, by paying people to rate a steak.
They posted a survey on Amazon Mechanical Turk, a site where people do small paid tasks, in 2016 and again in 2017. In all, 398 people from the United States finished it, for $1.75 each. Each rated 59 words from 0 to 100 on three scales: very working class to very upper class, very feminine to very masculine, and white to African American. The words came from seven domains, among them occupations, foods, clothing, sports and first names. The authors weighted the answers to match the national population by sex, education and race.
The projections and the survey averages agreed to a degree that depends on the dimension. For class the correlation was .53 on the Google Books model and .58 and .57 on two models trained on news and the open web. For gender it was .76, .88 and .90. For race it was .27, .75 and .44. A second test asked a humbler question: of any two words, does the model put them in the same order as the people? For class the Books model got 85 percent of pairs right, for gender 75 percent, for race 69 percent. The lesson is that every dimension has to earn its own validation. The gender ruler is sharp, the class ruler is usable, and the race ruler built from books deserves no trust on its own.
Others have checked the ruler against records rather than surveys. In 2018 Nikhil Garg and colleagues at Stanford measured, in a Google News embedding, how feminine each occupation word projected, and set that against the share of women in each occupation in the census. The two agreed with an r-squared of .46, and the fitted line ran through the origin, so an occupation split fifty-fifty showed no bias either way. A year earlier Aylin Caliskan, Joanna Bryson and Arvind Narayanan had found a correlation near .90 between the gender association of 50 occupation words and the Bureau of Labor Statistics share of women in them. Language carries the statistics of the society that wrote it, down to the payroll.
04Angles, error bars and timethe part the other methods cannot do
The ruler gives three things that word counts cannot: the angle between two distinctions, an error bar on every reading, and the same measurement repeated across decades.
Two dimensions are themselves arrows, so you can measure the angle between them. The authors found that cultivation and employment meet at 90 degrees, status sits at 85.5 degrees from cultivation and at 79.3 degrees from employment. No flat page can hold those three angles, since 85.5 plus 79.3 is more than 90, and that is the paper’s plain demonstration that class in American books is not one thing. For a clinic the same measurement would ask whether the blame distinction in notes leans toward the race distinction, the insurance distinction or neither.
The error bars come from a trick that any corpus allows. The authors cut the Google Books text into 20 pieces, trained a model on each, and read the projection off all 20. The band that holds 90 percent of the 20 readings is the error bar. A word that appears fewer than about 25 times wobbles so much that the band swallows the reading, and the paper drops such words from the map. A rare word is not a small signal. It is no signal.
Time comes from training one model per decade and laying the same ruler on each. Across the twentieth century the cosine between the education dimension and the affluence dimension climbed from about .15 in the 1900s to about .42 in the 1990s, while the other distinctions held their angles. In the books, schooling became what money means. The same design would serve a hospital whose notes span 20 years, with one model per year or per five-year block and the same antonym pairs each time.
05Five habits before you believe a numberthe ruler is yours, the reading is only as good as the ruler
Everything the method can get wrong follows from one fact: it measures how a community of writers used words, and nothing else.
First, whose corpus. Google Books is written by people who publish books, and a hospital’s notes are written by the people who write notes. The map is of the writers’ culture. Second, words, not people. If “nurse” projects feminine, that is a fact about print, and no nurse was asked. Third, the band. Report the 90 percent band with every reading, and do not read a word that falls below the frequency floor. Fourth, pairs first. The antonym pairs define the ruler, so write them down, and the reason for each, before looking at any projection, since a ruler chosen after the result will always confirm it. Fifth, flukes. A single reading that surprises you is a lead to check in a second corpus, not a finding. Alina Arseniev-Koehler’s 2024 review in Sociological Methods & Research walks through these choices one by one and is the best single thing to read after the paper itself.
06Three kinds of embeddingsword2vec is not the thing you send to OpenAI
The word “embedding” now names three different objects, and only the first one is a ruler.
Word2vec gives one fixed address per word type. “Bank” has one vector, parked halfway between the river and the money, and it is the same vector every time. That is a weakness for reading a sentence and a strength for measuring a culture, since the vector is a summary of every context the word ever had in that corpus.
A transformer, the kind of model inside a chatbot, gives a new vector for each use of a word, at each of many layers. The “bank” in “we sat on the river bank” and the “bank” in “the bank refused the loan” receive different vectors, because every layer rewrites the word in the light of its neighbors. These contextual embeddings are what make the model good at reading. They are harder to use as a ruler, because there is no single address for a word to project, though two recent papers from Kozlowski’s group show it can be done by averaging a word’s hidden states over several prompts.
When people “send text to OpenAI for embeddings,” they get a third thing. The embeddings endpoint takes a whole passage, a sentence or a page, and returns one vector of 1,536 or 3,072 numbers for the passage, not for any word in it. Their makers train them with a contrastive objective: pull together texts that belong together, a question and its answer, a passage and its paraphrase, and push unrelated texts apart. The geometry is tuned to answer one question well, which is “which of these passages are about the same thing,” the question search engines ask. The vendor does not disclose the training text, so nobody can date it or cut it into 20 pieces, and the vendor can replace the model at any time, after which the old numbers no longer compare with the new ones. Such vectors are excellent for finding the 50 notes most like this one. They cannot tell you what a word meant.
The two recent papers matter for one reason here. They show that the subtraction travels. Kozlowski, Callin Dai and Andrei Boutyline took the fixed word table at the front of Google’s Gemma models, built 28 antonym dimensions, projected 301 words, and found correlations with human ratings between .3 and .7. Kozlowski and Boutyline then did the same inside the hidden layers of Meta’s Llama models, for 360 words on 32 dimensions, checked against ratings from 1,750 people, and found the best dimensions above .8. They also found that the angles between the dimensions inside the model predict the correlations between the same scales in the human survey, with r of .87. The model’s geometry and the people’s judgments have the same shape. The interpretability people who steer chatbots by adding a direction to a hidden state are doing rich minus poor with a different corpus.
07Health, measured this way alreadydisease, stigma and the intensive care unit
The ruler has been laid on health five times already, in newspapers and in clinical notes, and each study shows one thing it can do.
Alina Arseniev-Koehler and Jacob Foster trained word2vec on the New York Times and built four rulers: gender, morality, health and class. They placed words about body weight on each. Fat projected immoral, unhealthy, poor and feminine, and the four readings moved separately. The newspaper carried four meanings of fat, and the method could say which was strongest.
Rachel Kahn Best and Arseniev-Koehler did the same for 106 diseases in 4.7 million news articles from 1980 to 2018. They built rulers for immorality, bad character and disgust. Preventable and behavioral conditions, such as addiction and obesity, sat at the immoral end. Infectious diseases sat at the disgust end. Over four decades stigma fell for chronic physical illness and for nothing else. One chart holds all 106 conditions, and that is the design to copy.
Julien Cobert and colleagues trained word2vec on intensive care notes from two hospitals, San Francisco from 2012 to 2022 and Boston from 2001 to 2012. They measured how close the words for Black and White patients sat to words for violence and noncompliance. In Boston the Black words sat closer to violence. In San Francisco the White words did. The notes of two hospitals carry two different cultures, and a model trained on either carries it too. Haoran Zhang and colleagues showed the same for BERT trained on the MIMIC-III notes: a model built on the notes performed differently by gender, language, ethnicity and insurance.
Two studies without embeddings show what the words in a note do. Anna Goddu and colleagues gave 413 physicians in training a chart note about a 28-year-old man with sickle cell disease. Half read a note that called him difficult, half read a neutral note about the same facts. The first half liked him less and treated his pain less aggressively. Michael Sun and colleagues searched 40,113 notes at one hospital for words such as refused, not adherent and combative. Black patients were two and a half times as likely to carry one. A note is a treatment the next clinician receives, and it is not given out evenly.
08When the corpus is smallyou do not need a hundred million words
Kozlowski had a century of books. Most health projects have a few thousand documents, and three tools make the ruler work at that size.
Facebook’s fastText project published static word vectors for 157 languages, Russian and Kazakh among them, trained on Wikipedia and a crawl of the web, 300 numbers per word, built from pieces of words so that a noun with twelve suffixes shares its evidence across all twelve forms. For biomedical English there is BioWordVec, published in Scientific Data in 2019 and trained on PubMed abstracts and the MIMIC-III notes. Pretrained vectors carry the full Kozlowski recipe, antonym pairs and all, with no training at all. What they cannot do is tell you how your community used a word, because the map is someone else’s.
Pedro Rodriguez, Arthur Spirling and Brandon Stewart solved that in 2023 with à la carte embeddings. For every mention of your word in your corpus, average the pretrained vectors of the words around it, then multiply the average by a matrix learned once from a large corpus to undo the blurring. The result is a vector for your word in your text, and it works with fewer than twenty mentions. Their R package, conText, compares the vector between groups, such as notes from two departments or two years, with a permutation test for the difference.
Dustin Stoltz and Marshall Taylor approached from the other end with concept mover’s distance in 2019. Instead of asking where a word sits, it asks how far the words of one document would have to travel across the map to reach a concept, such as death or blame. A short trip means the document engages the concept. It works on a tweet or a novel, with pretrained vectors, and it gives a score per document, which is the natural unit when the documents are notes.
09A note from Claude on patient recordspain is the best case for the ruler
Dina asked me what the ruler could do with health data. This section is my answer, and it is signed.
In a hospital the ruler measures the clinicians. Notes are written by clinicians, so a map built from notes is a map of how clinicians write about patients. That is the opportunity. Goddu showed that the words in a note change what the next reader does, and Sun counted the words. The ruler adds the angle. Build a blame ruler from pairs such as compliant and noncompliant, cooperative and combative. Build an insurance ruler and a race ruler from the words the notes use. Measure the angle between blame and each. If blame leans toward the uninsured pole and stands at a right angle to race, the notes have said where their moral vocabulary lives. Put 100 diagnoses on the blame ruler and the chart shows which conditions carry it.
The number belongs to a word, never to a patient. A projection says how a community of writers used a word. It says nothing about a person in the record, and it must not be turned into a score for one. That sentence goes into the ethics application, and it is also the method’s best defense. Nothing in the output identifies, ranks or predicts anyone.
Pain is the best case for the ruler, because pain exists in medicine only as words. A blood sugar is a number from a machine. Pain is a number the patient picks from 0 to 10 and the words the patient and the clinician use around it. Those words are the data, and the ruler is a method for words. Three measurements follow.
First, a pain-severity ruler. The McGill Pain Questionnaire, written by Ronald Melzack in 1975, is a list of 78 pain words, such as throbbing, stabbing, gnawing and unbearable, sorted into 20 groups and ranked by intensity from human ratings. Those ranks are a ready-made human panel, like Kozlowski’s 398 steak raters. Build a severity ruler from pairs such as mild and severe, bearable and unbearable, and check the projections of the 78 words against Melzack’s ranks. A ruler that passes can score any pain vocabulary at all, including words the questionnaire never listed and words in Russian or Kazakh, where no ranked pain word list exists.
Second, a credibility ruler. Kelly Hoffman and colleagues showed in 2016 that half of a sample of white medical students and residents held false beliefs about Black patients’ bodies and rated their pain lower. In notes that bias has a vocabulary. Patients report, state and endorse pain, or they claim, insist and complain of it. Build the ruler from believed and doubted, reports and claims, and project the words that surround pain in notes by patient group, by department and by year. The 2016 opioid prescribing guideline is a before and after. If pain words moved toward the doubted end after it, the notes show the guideline’s effect on language, with an error band.
Third, the translation gap. A patient writes that her chest is being crushed. The note says chest pressure. Project patient words and chart words for the same complaint on the same severity ruler, and the distance between them measures what the chart lost. Portal messages, complaints and forum posts give the patient side. The same test fits chronic, stable and routine, which clinicians use to reassure and patients hear otherwise.
Notes need cleaning before training, and the cleaning is most of the work. Templates repeat ten thousand times and pull every word in them together, so strip them. Negation puts chest pain next to denies, so mark negated spans, or the model learns that pain is a thing patients deny. Abbreviations need a pass, or SOB lands somewhere strange. Check the nearest neighbors of twenty common symptoms by hand before believing anything.
Validate with two panels. Have clinicians and patients rate the same words on the same scales. The disagreement between the two panels is a finding on its own. Sun’s hand-coded word list is the other test. A stigma ruler that predicts the hand codes has earned trust.
Vectors are patient data. John Morris and colleagues recovered full names from clinical note embeddings in 2023. Train inside the institution, publish projections and bands and never the vectors of rare words, and send nothing to a vendor’s endpoint. The frequency floor the method needs anyway, dropping words seen fewer than 25 times, is also a privacy floor.
Write the pairs down first. The antonym pairs are the hypothesis. A team that registers its pain pairs, blame pairs and social pairs before seeing a projection has a study. A team that picks pairs while looking has a mirror. The first deliverable is one page: the rulers, the pairs behind each, the corpora, and what a null result looks like.
Claude
Anthropic’s model, writing with Dina Pisareva, 7 October 2026
10Reading listin the order the page uses them
- Kozlowski, Taddy and Evans, The Geometry of Culture: Analyzing the Meanings of Class through Word Embeddings, American Sociological Review 84(5), 2019
- Kozlowski and Boutyline, The Semantic Structure of Feature Space in Large Language Models, arXiv 2604.27169, 2026
- Kozlowski, Dai and Boutyline, Semantic Structure in Large Language Model Embeddings, arXiv 2508.10003, 2025
- Arseniev-Koehler, Theoretical Foundations and Limits of Word Embeddings: What Types of Meaning Can They Capture?, Sociological Methods & Research 53(4), 2024
- Arseniev-Koehler and Foster, Machine Learning as a Model for Cultural Learning: Teaching an Algorithm What It Means to Be Fat, Sociological Methods & Research 51(4), 2022
- Best and Arseniev-Koehler, The Stigma of Diseases: Unequal Burden, Uneven Decline, American Sociological Review 88(5), 2023
- Cobert et al., Measuring Implicit Bias in ICU Notes Using Word-Embedding Neural Network Models, Chest 165(6), 2024
- Zhang, Lu, Abdalla, McDermott and Ghassemi, Hurtful Words: Quantifying Biases in Clinical Contextual Word Embeddings, ACM CHIL, 2020
- Goddu et al., Do Words Matter? Stigmatizing Language and the Transmission of Bias in the Medical Record, Journal of General Internal Medicine 33(5), 2018
- Sun, Oliwa, Peek and Tung, Negative Patient Descriptors: Documenting Racial Bias in the Electronic Health Record, Health Affairs 41(2), 2022
- Garg, Schiebinger, Jurafsky and Zou, Word Embeddings Quantify 100 Years of Gender and Ethnic Stereotypes, PNAS 115(16), 2018
- Caliskan, Bryson and Narayanan, Semantics Derived Automatically from Language Corpora Contain Human-like Biases, Science 356, 2017
- Charlesworth, Caliskan and Banaji, Historical Representations of Social Groups across 200 Years of Word Embeddings from Google Books, PNAS 119(28), 2022
- Rodriguez, Spirling and Stewart, Embedding Regression: Models for Context-Specific Description and Inference, American Political Science Review 117(4), 2023
- Stoltz and Taylor, Concept Mover’s Distance, Journal of Computational Social Science 2, 2019
- Zhang, Chen, Yang, Lin and Lu, BioWordVec, Improving Biomedical Word Embeddings with Subword Information and MeSH, Scientific Data 6, 2019
- Morris, Kuleshov, Shmatikov and Rush, Text Embeddings Reveal (Almost) As Much As Text, EMNLP, 2023
- Melzack, The McGill Pain Questionnaire: Major Properties and Scoring Methods, Pain 1(3), 1975
- Hoffman, Trawalter, Axt and Oliver, Racial Bias in Pain Assessment and Treatment Recommendations, PNAS 113(16), 2016
- Mikolov, Chen, Corrado and Dean, Efficient Estimation of Word Representations in Vector Space, 2013
- Grave, Bojanowski, Gupta, Joulin and Mikolov, Learning Word Vectors for 157 Languages, LREC, 2018