Unilanguage Experimental Research · UNI-EXP-001A-R2

M → HUMAN Semantic Association

Can a proposed cross-language pattern survive a predefined empirical test?

M → HUMAN 语义关联:一个跨语言观察,能否经受预先确定的实证检验?

Preregistered · Round 2 Rules frozen Exploratory AI result published Human annotation pending

Exploratory AI Result

Inconclusive—human result still pending

Public-safe record

Current status

  • AI-A: M 14.90% vs Controls 11.76%; risk difference +3.14 pp
  • AI-B: M 15.49% vs Controls 11.96%; risk difference +3.53 pp
  • Both 95% confidence intervals include zero
  • Evidence-family analysis preserves the positive direction
  • Pre-unblinding methodological review archived publicly · 揭盲前方法审查已公开归档

Research boundary

Exploratory AI Result — Inconclusive. Both frozen passes point in the predicted direction, but every primary and evidence-family 95% confidence interval includes zero and every directional Fisher test has p > 0.05. This result does not replace or decide the preregistered human blind analysis.

探索性 AI 结果——无法判断。两份冻结标注均指向预测方向,但所有主要分析和证据家族分析的 95% 置信区间都包含 0,且所有单侧 Fisher 检验均为 p > 0.05。本结果不替代、也不裁定预注册真人盲标分析。

Open the Public-safe AI Annotation Record · 查看公开安全记录 Read the Historical Blinded Methodological Review · 查看揭盲前历史方法审查

Research question

Does initial M predict HUMAN meaning?

This experiment tests whether English words beginning orthographically with M are disproportionately associated with the semantic domain HUMAN, compared with predefined control initial-letter groups. It tests a lexical distributional tendency—not shared etymology, universal sound symbolism, or causation.

本实验检验:英语中以字母 M 开头的词,是否相对于预先确定的对照首字母组,更集中地分布于 HUMAN(人类)语义场。本实验只检验词汇分布倾向,不证明共同词源、普遍语音象征或因果关系。

Pilot study

Pilot summary

Completed

What the Pilot established

The Pilot made the M → HUMAN proposal explicit and showed that it could be converted into a testable comparison. It also exposed major sources of researcher freedom: hand-picked examples, unstable semantic boundaries, related-word duplication, and control selection.

Pilot 将 M → HUMAN 观察转化为可比较的假说,同时暴露了人工选词、语义边界不稳定、同族词重复计数和对照选择等研究者自由度。

Why a second round is necessary

The Pilot is exploratory evidence only. It cannot decide the hypothesis. Round 2 therefore freezes definitions, sampling, annotation, counterexample handling, and reporting rules before outcomes are inspected.

Pilot 只提供探索性证据,不能裁决假说。Round 2 因而在查看结果前冻结定义、抽样、标注、反例处理和报告规则。

Round 2

Preregistration status

Frozen · v1.0
Version 1.019 August 2026Author: Jinkai Liu

Primary rules are fixed before outcome inspection.

The 31-section preregistration defines the claim, categories, sampling principle, controls, frequency matching, evidence-family deduplication, blind coding, analyses, reporting standard, and decision rule.

31 节预注册文件已经固定假说范围、语义类别、抽样原则、对照组、频率匹配、证据家族去重、盲标、分析、报告和结论规则。

View Full 31-Section Preregistration | 查看完整31项预注册规则

Immutable public record · 不可覆盖的公开记录: Version v1.0 · 2026-08-19 · Jinkai Liu · UNI-EXP-001A-R2 · Frozen / Preregistered. Future framework versions do not replace this record.

Formal hypotheses

H₀ and H₁

H₀ · Null hypothesis

P(HUMAN | M) ≤ P(HUMAN | Controls)

M-initial words are not more likely to belong to HUMAN than words in the predefined controls.

M 开头的词进入 HUMAN 语义场的比例不高于预先设定的对照组。

H₁ · Research hypothesis

P(HUMAN | M) > P(HUMAN | Controls)

M-initial words are more likely to belong to HUMAN than words in the predefined controls.

M 开头的词进入 HUMAN 语义场的比例高于预先设定的对照组。

Method

Sampling and semantic coding

01

Frequency-based sampling

Eligible modern English lexical items are sampled from a predefined frequency source. The M group and pooled controls are matched as closely as practical across fixed frequency bands.

从预先确定的现代英语词频来源抽样,并在固定频段内尽可能匹配 M 组与合并对照组。

02

Evidence-family control

Related forms receive an Evidence Family ID. Results are reported both by eligible lexeme and after related forms are collapsed.

相关词形获得 Evidence Family ID;结果同时报告词项层分析与证据家族去重分析。

03

Primary-sense rule

HUMAN coding uses the primary or highest-frequency contemporary dictionary sense. Rare or historical senses cannot rescue a favorable case.

HUMAN 判断采用当代词典的第一义或最高频义;罕见义和古义不能用于事后挽救支持案例。

Bias control

Blind annotation

What annotators see

  • Lexical item
  • Standardized dictionary definition
  • Fixed semantic coding instructions
  • HUMAN, PERSON, PEOPLE, IDENTITY, HUMAN-ATTRIBUTE, and UNCERTAIN labels

标注者只看到词项、标准词典释义、固定分类说明和语义标签。

What annotators do not see

  • The “M hypothesis” label
  • The expected direction
  • Pilot outcomes
  • Lists of supposed supporting words

标注界面不出现 M 假说名称、预期方向、Pilot 结果或所谓支持词表。

Two annotators code each item independently. Initial disagreements, agreement statistics, adjudication, and both original labels are preserved.

每个词项由两名标注者独立编码;原始分歧、一致性统计、裁决结果及两份原标签全部保留。

Counterexamples policy

Non-supportive cases are data.

Every eligible M-initial word remains in the dataset regardless of category. Words such as machine, metal, mountain, music, and minute, if sampled, cannot be removed because they weaken the proposal.

所有符合资格的 M 开头词都必须留在数据中。若 machinemetalmountainmusicminute 等词进入样本,不能因其削弱假说而被删除。

Supporting cases · Counterexamples · Neutral cases · Excluded cases with reasons

最终数据同时保存:支持案例、反例、中性案例,以及附排除理由的排除案例。

Reporting standard

Effect size and uncertainty first

Risk differenceAbsolute difference in HUMAN proportions · HUMAN 比例绝对差
Odds ratioRelative association measure · 相对关联指标
95% CIPrecision and plausible effect range · 精度及合理效应范围
p-value + NReported, but never used alone · 报告但不单独裁决假说

The interpretation prioritizes effect direction, magnitude, confidence interval, robustness after evidence-family deduplication, and then statistical significance.

解释顺序优先考虑效应方向、效应大小、置信区间、证据家族去重后的稳健性,最后才是统计显著性。

Round 2 outcomes

Exploratory AI Result

Inconclusive

Both frozen AI passes show a modest positive M-minus-Control HUMAN difference, while both 95% confidence intervals include zero. The evidence-family robustness analysis retains the same direction but remains imprecise. Under the requested three-level vocabulary, the result is Inconclusive. This is an Exploratory AI Result; Preregistered Human Annotation: Pending.

两份冻结 AI 标注均显示 M 组 HUMAN 比例略高于对照组,但两个 95% 置信区间都包含 0;证据家族去重分析保持同一方向,但精度仍不足。按本次要求的三级结论词汇,本结果为无法判断。本结果仅为探索性 AI 结果预注册真人标注:待进行。

AI-A: 14.90% vs 11.76%76/510 M vs 60/510 Controls · M 组与对照组
+3.14 pp95% CI −1.05 to +7.33 · one-sided Fisher p=0.0835
AI-B: 15.49% vs 11.96%79/510 M vs 61/510 Controls · M 组与对照组
+3.53 pp95% CI −0.71 to +7.77 · one-sided Fisher p=0.0608
Read the complete exploratory AI record · 查看完整探索性 AI 记录 Read the bilingual conclusion clarification · 查看双语结论澄清

Decision record

Human assessment remains pending

Preregistered human result pending

Exploratory AI Result — Inconclusive: the estimated HUMAN effect is positive and consistent across AI-A, AI-B, and fixed-family robustness analyses, but all uncertainty intervals include zero and all directional Fisher tests exceed 0.05. The archived v1.0 five-level wording “Tentatively Supported” is preserved as a historical record; this three-level classification is the current public-safe conclusion. No preregistered human status has been assigned.

探索性 AI 结果——无法判断:AI-A、AI-B 与冻结证据家族稳健性分析均给出正向 HUMAN 效应,但所有不确定性区间都包含 0,且所有单侧 Fisher 检验均超过 0.05。v1.0 采用五级词汇的“初步支持”作为历史记录原样保留;当前公开安全的三级结论为“无法判断”。预注册真人结论尚未分配。

We publish hypotheses—including hypotheses our own experiments fail to support.

我们公开记录假说,也公开记录实验未能支持的假说。