Balancing Accuracy and Efficiency: Evaluating Encoder- and Decoder-Based Models for Word Sense Disambiguation and Regular Polysemy Detection
This repository contains the code accompanying the article "Balancing Accuracy and Efficiency: Evaluating Encoder- and Decoder-Based Models for Word Sense Disambiguation and Regular Polysemy Detection" by Pin-Er Chen, Da-Chen Lian, and Shu-Kai Hsieh (Graduate Institute of Linguistics, National Taiwan University), published in Natural Language Processing 32(4), 451–469 (Cambridge University Press). The version of record is available at Cambridge Core.
Publication: Chen, Lian, and Hsieh (2026), Natural Language Processing 32(4), 451–469. https://doi.org/10.1017/nlp.2026.10033
This study investigates the nuanced challenges of fine-grained Word Sense Disambiguation (WSD) tasks with Regular Polysemy Detection (RPD) of the Named Entity, focusing on evaluating the trade-offs between encoder and decoder-based model performance and computational efficiency. The datasets, including Chinese Wordnet 2.0 (CWN) as sense inventory, the Social Media Corpus (PTT) for user-generated content, and the Academia Sinica Balanced Corpus (ASBC) for formal linguistic data, were chosen to provide a diverse and representative framework for evaluating both common nouns and proper nouns with regular polysemy in Taiwan Mandarin. This analysis evaluated ten encoder- and decoder-based models, assessing their performance on two tasks. The encoder-based models demonstrate comparable accuracy to the decoder-based models on WSD tasks (77.5 per cent vs. 78.5 per cent), and similarly strong performance in RPD tasks (84.2 per cent vs. 83.8 per cent). On a large-scale all-words word sense disambiguation task, the encoder model not only outperformed the decoder model but also generated substantially lower carbon emissions—an eight-fold reduction. These differences underscore the trade-offs between model architecture and task-specific performance, highlighting the necessity for balancing performance and energy efficiency in the design and application of language models, advocating for sustainable and eco-friendly practices in NLP development.
Plain text:
Chen, P.-E., Lian, D.-C., & Hsieh, S.-K. (2026). Balancing accuracy and efficiency: Evaluating encoder- and decoder-based models for word sense disambiguation and regular polysemy detection. Natural Language Processing, 32(4), 451–469. https://doi.org/10.1017/nlp.2026.10033
BibTeX:
@article{Chen_Lian_Hsieh_2026,
title = {Balancing accuracy and efficiency: Evaluating encoder- and decoder-based models for word sense disambiguation and regular polysemy detection},
volume = {32},
doi = {10.1017/nlp.2026.10033},
number = {4},
journal = {Natural Language Processing},
author = {Chen, Pin-Er and Lian, Da-Chen and Hsieh, Shu-Kai},
year = {2026},
pages = {451--469}
}See CITATION.cff for the machine-readable form.
The tagging layer in this repository was derived from the experimental
dwsd-beta
implementation in lopentu/CwnSenseTagger. The link is pinned to commit
db384a3f8bc5da7f227c0f46fb515ddf8056dbcd, where the relevant upstream
implementation was last changed; linking to the mutable master branch would
not identify a reproducible source version.
This repository is not an unmodified copy of, or a runtime dependency on, the
CwnSenseTagger Python package. It preserves the original candidate-generation
and decoding workflow while generalizing the model, training, and evaluation
infrastructure used for the experiments in the accompanying article. The
derived implementation has been substantially modified for this project since
2024; the material changes are summarized below.
| Area | Preserved from dwsd-beta |
Changed in this repository |
|---|---|---|
| Target and candidate construction | Marks the target as <word> and pairs its context with word, sense definition, first reference example |
Adds typed training, testing, and inference representations; the actual inference text construction is unchanged |
| CWN candidate selection | Looks up all senses, filters by compatible POS, and excludes senses without examples | Increases the candidate-lookup cache from 1,000 to 10,000 entries |
| Regular polysemy detection | Falls back to RPD for Nb/Nc tokens without WSD candidates and uses the same RP glossary (glossdict.json is byte-identical) |
Adds explicit RPD datasets, evaluation paths, and classification reports |
| Prediction decoding | Applies softmax across the candidates for each example and selects the highest-probability candidate | Adds typed prediction structures and support for Hugging Face model output conventions |
| Model | Uses a single bert-base-chinese model with BERT pooler_output and a binary classification head |
Uses AutoModelForSequenceClassification and each base model's native classification behavior for the reported experiments; also retains an AutoModel compatibility fallback for custom state dicts |
| Checkpoint loading | Downloads one fixed checkpoint from Google Drive | Loads released Hugging Face checkpoints or an explicitly supplied local state dict; no automatic Google Drive download |
| Chinese Wordnet data | Uses CwnImage.latest(), whose result can change as new images become available |
Loads the explicit CWN image v.2022.08.01 for reproducibility |
| Batching and tokenization | Uses paired context/candidate inputs, maximum length 320, and inference batch size 16 | Makes batching configurable, pads to a multiple of 8 instead of 16, supports decoder padding, and can remove unsupported token_type_ids |
| Experimental scope | Provides the original single-model inference tagger | Adds multi-model fine-tuning, LoRA, held-out WSD/RPD evaluation, ASBC large-scale evaluation, and carbon-emissions analysis |
The models reported in the article were trained and evaluated through
AutoModelForSequenceClassification; their native sequence-classification
heads produce the logits. They do not use the explicit mean pooling in
DottedWsd.
DottedWsd is a compatibility path for loading a base AutoModel plus a
separate local state dict when a sequence-classification checkpoint is not
available. It reduces token-level hidden states to one vector by taking their
mean. This avoids assuming that every architecture supplies BERT's
pooler_output, but the current fallback does not mask padding positions.
Anyone extending this path to padded batches should replace the plain mean with
attention-mask-aware pooling. The supplied training and evaluation scripts do
not select this fallback.
The direct lexical-resource dependency is
CwnGraph, declared as cwngraph>=0.4.0;
the committed lockfile resolves it to version 0.4.0. The relevant local
implementations are in src/dotted_wsd/tagger/ and
src/dotted_wsd/dwsd_datasets.py.
| Path | Purpose |
|---|---|
src/dotted_wsd/ |
Library code: dataset loaders, training, evaluation, ASBC large-scale tagging |
src/dotted_wsd/asbc_eval/ |
§5 large-scale ASBC tagging pipeline (see its own README.md) |
scripts/ |
Shell entry points for training, held-out evaluation, and ASBC large-scale tagging |
tokenizers/{configs,customized,default}/ |
Per-base tokenizer YAMLs and their customized / reference JSONs |
tokenizers/customize.py |
Reproduces tokenizers/customized/*.json from tokenizers/default/*.json |
carbon_emissions/ |
Reproduction package for the paper's emissions numbers (see its own README.md) |
tests/ |
Pytest suite: imports, CLI --help, carbon_emissions reproducibility, deduplicate_instances round-trip, customize.py byte-equality |
pyproject.toml / uv.lock |
uv-managed Python environment |
This project uses uv for environment management.
# Install uv if you don't already have it
curl -LsSf https://astral.sh/uv/install.sh | sh
# Sync the locked environment (creates .venv/, installs all pinned dependencies)
uv syncRequires Python ≥ 3.10. PyTorch is pulled from the CUDA 12.1 wheel index (see
[tool.uv.sources] in pyproject.toml); machines without a CUDA 12.1-compatible
GPU may need to adjust the index before running uv sync.
uv.lock is committed and pins the dependency versions this release was
tested against. Note that the lock has been refreshed since the original
internal training environment (e.g. to add pandas/numpy/tqdm as
direct deps and to bump a yanked transitive protobuf version), so it
is not byte-identical to the snapshot that produced the paper's released
weights, but resolved versions of the load-bearing libraries
(torch, transformers, peft, accelerate, etc.) are unchanged.
No data ships with this repository. You must obtain and place the
following files under a top-level data/ directory yourself. The
data/.gitkeep placeholder is committed only so the directory exists
on a fresh clone.
data/
├── WSD_merge_train_v2.csv # WSD training split (context-gloss pairs, "Reproducing training")
├── WSD_merge_test_v2.csv # WSD held-out test split (used by training and evaluation)
├── RP_train.csv # RPD training split
├── RP_valid.csv # RPD held-out test split
├── wsd_examples.csv # By-example WSD eval set (used by `dwsd_eval`)
├── glossdict.json # CWN gloss dictionary
└── dt-asbc/ # Raw ASBC tagged .txt files (input to §5 preprocessing)
├── asbc_dotted_tagged_000-of-140.txt
├── asbc_dotted_tagged_001-of-140.txt
└── ... # 140 files in total
Two further files are produced by the preprocessing pipeline (see the Preprocessing section), not supplied:
data/dt_asbc_dataset/*.csv— per-file WSD instance CSVs fromprocess_into_instances.data/asbc-deduplicated-instances.csv— the deduplicated dataset, human-readable.data/asbc-deduplicated-instances.feather— the same dataset as zstd-compressed feather; this is whatscripts/run_asbc_eval.shreads.
The data files listed above are not produced by this repo and are not shipped with it.
The library reads these paths via dotted_wsd.settings.DATA_DIR, which
defaults to <repo>/data/.
The 140 raw ASBC tagged .txt files are turned into the deduplicated
feather dataset that scripts/run_asbc_eval.sh consumes via two scripts.
Set up the input directory once, then run the two commands.
-
Place raw files under
data/dt-asbc/— see the Data section. -
Per-file WSD instance extraction (looks up candidate senses in CWN):
uv run python -m dotted_wsd.asbc_eval.process_into_instances
Outputs one
*-eval-prepared.csvper input file underdata/dt_asbc_dataset/(the directory is created automatically). Add--debugto process only the first two input files for a smoke test. -
Cross-file deduplication of
(test_sentence, test_word)pairs:uv run python -m dotted_wsd.asbc_eval.deduplicate_instances
Outputs three files at the top of
data/:asbc-deduplicated-instances.csv— human-readableasbc-deduplicated-instances.feather— zstd-compressed; whatscripts/run_asbc_eval.shreadstest_sentence_to_example_ids.json— sidecar mapping each deduplicatedtest_sentenceback to everyexample_idthat was collapsed into it
Pass
--save-dir <path>to redirect.
The decoder-based bases (Llama-3.2-3B, Gemma-2-2b, SmolLM-180M) use
customized tokenizer JSONs vendored under tokenizers/customized/; the
unmodified reference JSONs are at tokenizers/default/. The customization
registers one or two special tokens with each tokenizer's Rust-level
post_processor and rewrites the pair template; per-base details are
encoded as deterministic patch functions in
tokenizers/customize.py:
uv run python tokenizers/customize.py # rewrite tokenizers/customized/*.json
uv run python tokenizers/customize.py --check # diff against vendored files; exit 1 on driftTraining and evaluation scripts read the customized files via the per-base
YAMLs in tokenizers/configs/. The shipped JSONs are sufficient for
reproduction — no extra step. Run customize.py only if you want to adapt
the customization to a new base model (use one of the existing patch
functions as a template).
The training data CSVs are not produced by this repo and are not openly available — see the Data section.
13 lopentu/*-DottedWSD checkpoints are released under the
lopentu organization and are publicly
available. The paper itself reports on a subset of 10 of these.
lopentu/google-bert-bert-base-chinese-DottedWSDlopentu/ckiplab-bert-base-chinese-DottedWSDlopentu/yentinglin-bert-base-zhtw-DottedWSDlopentu/IDEA-CCNL-Erlangshen-DeBERTa-v2-97M-Chinese-DottedWSDlopentu/MoritzLaurer-mDeBERTa-v3-base-xnli-multilingual-nli-2mil7-DottedWSDlopentu/MoritzLaurer-mDeBERTa-v3-base-mnli-xnli-DottedWSDlopentu/microsoft-mdeberta-v3-base-DottedWSDlopentu/microsoft-deberta-v3-small-DottedWSDlopentu/microsoft-deberta-v3-base-DottedWSDlopentu/microsoft-deberta-v3-large-DottedWSDlopentu/SmolLM-Chinese-180M-DottedWSDlopentu/gemma-2-2b-DottedWSDlopentu/meta-llama-Llama-3.2-3B-DottedWSD
The training entry point is dotted_wsd.train.hf_trainer. To reproduce the
full grid of fine-tunes from the paper:
# Edit scripts/run_hf_trainer.sh to comment out any configs you don't want
bash scripts/run_hf_trainer.shBy default:
- Training metrics are not reported anywhere (
--report-to none). - The trained model is not pushed to the Hugging Face Hub (
--no-push-to-hub).
To log to Weights & Biases, pass --report-to wandb and (optionally) override
the destination via env vars:
WANDB_ENTITY=your-entity WANDB_PROJECT=your-project \
uv run python -m dotted_wsd.train.hf_trainer <model_id> <batch_size> --report-to wandbTo push the trained model to the Hub, pass --push-to-hub and a
--hub-model-id you can write to:
uv run python -m dotted_wsd.train.hf_trainer <model_id> <batch_size> \
--push-to-hub --hub-model-id your-namespace/your-model-nameIf --hub-model-id is omitted, it defaults to
lopentu/{model_id}-DottedWSD{suffix} — only useful if you have write access
to the lopentu org.
bash scripts/run_eval.shReads each released lopentu/*-DottedWSD model from the Hugging Face Hub
and writes per-model evaluation results under
data/eval_results/<model_name>/ (gitignored). The output for each model
is two pickle files (*_wsd_eval.pkl and *_rp_eval.pkl) plus a few
PNG plots from the by-example analysis. The WSD pickle contains
{metadata, by_example, by_instance} (per-example/instance prediction
DataFrames + accuracy values); the RP pickle contains
{hint, nohint}, each holding a sklearn-style classification report
DataFrame and a per-row prediction list.
The §5 experiment runs each fine-tuned model over the full ASBC corpus
(~32.7M instances). It assumes you've already produced
data/asbc-deduplicated-instances.feather via the
ASBC corpus preparation
preprocessing pipeline.
bash scripts/run_asbc_eval.shTagging outputs land under data/asbc_eval_results/<model_name>/. The
shipped run_asbc_eval.sh includes --debug for safety; remove the
--debug flag for a full production run. When using --debug with
the default --preprocess-workers (≈ cpu_count()), the tiny debug
dataset can trigger BrokenProcessPool; pass --preprocess-workers 8
in that case.
For the per-model energy and CO₂e numbers cited in the paper's abstract,
see carbon_emissions/ — input CSV plus three
stdlib-only Python scripts that reproduce the totals, run the
sensitivity analysis, and verify each headline number.
See src/dotted_wsd/asbc_eval/README.md
for a tighter end-to-end recipe.
This repository's source code is released under
GPL-3.0-only — see
LICENSE. This license preserves compatibility with the
GPL-3.0-licensed
CwnSenseTagger source from which the tagging layer was derived.
The released lopentu/*-DottedWSD model
checkpoints are not covered by this repository's source-code license. Each
checkpoint remains subject to the applicable license or terms of its upstream
base model, as identified on the checkpoint and base-model cards. The same
principle applies to datasets and other third-party artifacts: their original
terms continue to apply, and this repository does not relicense them.