Balanced sampling via the Cube Method (Deville & Tillé, 2004) applied to etymology — what share of the French lexicon carries an Arabic root, visible or masked by a Romance intermediary?
A balanced sampling algorithm that guarantees exact Horvitz-Thompson totals on all auxiliary variables, for any population.
Random walk in the kernel of the balancing matrix A. At each step, probabilities shift until one reaches 0 or 1 — preserving balancing constraints exactly throughout.
Rounds residual fractional probabilities with minimal perturbation. Bousabaa, Lieber & Sirolli (INSEE 1999) proved the existence of such balanced roundings.
Each unit k contributes yk / πk. The estimator is unbiased under any unequal-probability sampling design.
The design satisfies Âx ≈ Ax: estimated totals of auxiliary variables equal their true values, reducing variance vs. simple random sampling.
The Cube Method applies to any problem requiring a representative sample of a population with balancing constraints.
Employment, consumption, partial census surveys (INSEE, Eurostat)
Balanced randomisation over age, sex, comorbidities across groups
Transaction selection balanced on amount, type, entity
Corpus balanced on frequency, POS, semantic domain — this project
Forest inventory balanced on species, altitude, aspect
Panel balanced on income, age, region, education level
Enter your data (one unit per line: y z1 z2 …), choose sample size n and seed. The Cube algorithm draws a balanced subsample and computes the HT estimate of total(y).
Foundational paper. Efficient balanced sampling: the Cube method. Biometrika 91(4), 893-912.
Existence of balanced roundings. Working paper INSEE / CREST.
Sampling Algorithms. Springer. Comprehensive reference on balanced designs and HT estimation.
The Cube Method is applied to a corpus of French lemmas. The Horvitz-Thompson estimator quantifies the share of the vocabulary with Arabic etymology, direct or masked by a Romance intermediary.
TLFi, Larousse and Robert cite the proximal origin (Italian, Spanish, Medieval Latin) rather than the ultimate origin. jupe is listed as Italian, but the full chain is: Arabic jubba → Spanish → Italian → French.
A masked Arabic word is one where Arabic does not appear at the first position of the etymological chain cited by standard dictionaries, but does appear in the full chain traced to the ultimate origin.
566 French lemmas (nouns + adjectives) balanced on 5 auxiliary variables: log₁₀ frequency rank, part of speech, semantic domain, word length, al- prefix indicator.
Each word k contributes with weight 1/πk. The 95% confidence interval uses the Deville-Tillé variance under balanced design.
Manually curated from TLFi, CNRTL, Corriente (2008), Devic (1876). Data from etymology_database.py.
Source cited by standard dictionaries vs. full reconstructed chain.
| French word | Cited origin (TLFi) | Full chain | Arabic |
|---|---|---|---|
| jupe | Old French / Italian | Arabic jubba → Sp. aljuba → It. giuppa → Fr. jupe | masked |
| artichaut | Spanish / Italian | Ar. al-kharshuf → Sp. alcachofa → It. articiocco → Fr. | masked |
| café | Italian | Ar. qahwa → Turkish kahve → It. caffè → Fr. café | masked |
| arsenal | Italian | Ar. dār aṣ-ṣināʿa → It. arsenale → Fr. | masked |
| assassin | Italian | Ar. ḥashshāshīn → Med. Lat. → It. assassino → Fr. | masked |
| luth | Spanish / Provençal | Ar. al-ʿūd → Sp. laúd → It. liuto → Fr. luth | masked |
| coton | Italian | Ar. quṭun → It. cotone → Fr. coton | masked |
| douane | Spanish / Turkish | Ar. dīwān → Turkish → It. dogana → Fr. | masked |
| algèbre | Medieval Latin | Ar. al-jabr (al-Khwarizmi) → Med. Lat. algebra → Fr. | direct |
| algorithme | Medieval Latin | Ar. al-Khwarizmi (latinised proper name) → Lat. algorismus → Fr. | direct |
The Cube algorithm draws a balanced sample from the corpus, then queries Claude (Sonnet) word by word to trace full etymological chains.
Unit test result will appear here…