MACHINE LEARNING · EMPIRICAL
Original research: Using Codebooks to Detect Cybercrime Topics in Text Narratives · 2609.16000v1
Paper authors: Shufan Chai, Liangliang Sun, Jessica Staddon
Source license: CC BY-SA 4.0. This article summarizes and interprets the source using AI. Attribution does not imply endorsement by the source authors.
This adapted analysis is shared under the same CC BY-SA 4.0 license. Semantic status: supported by automated evidence review. Human scientific review and independent replication have not been established.
TL;DR
The proposed contribution is a generalist-oriented prompting method that uses qualitative codebooks to detect cybercrime incidents in consumer narratives with pretrained language models.
For impostor scams, the codebook condition achieved strong average precision and recall and improved average precision relative to the baseline.
Source: E010
For identity theft, baseline precision was weaker, while the codebook condition attained stronger average precision and recall and improved average precision over baseline.
Source: E006
In the reported case studies, every evaluated model apart from the identified exception surpassed the stated precision-and-recall threshold under codebook prompting.
Generalization may be limited to cybercrime attributes or events that are amenable to qualitative analysis and codebook development.
Source: E005
The study evaluates only a single prompt template and does not establish whether more elaborate templates would improve performance.
Source: E016
The findings suggest that resource-constrained organizations may be able to use pretrained models, potentially including lower-cost options, without specialized model development or domain expertise.
Source: E016
Significance
The proposed contribution is a generalist-oriented prompting method that uses qualitative codebooks to detect cybercrime incidents in consumer narratives with pretrained language models.
The findings suggest that resource-constrained organizations may be able to use pretrained models, potentially including lower-cost options, without specialized model development or domain expertise.
Source: E016
Research Question
The research examines whether general-purpose pretrained language models can identify cybercrime in narrative text when guided by researcher-authored coding guidance.
Source: E004
Contribution
Tested Scope
The empirical scope covers impostor scams and identity theft as the studied detection topics.
Source: E014
The approach assumes that its codebooks have been rigorously developed so that independent human annotators can apply them reliably and consistently.
Source: E001
Datasets
The impostor-scam corpus comprises deduplicated complaint narratives and includes both positive and negative labels.
Source: E017
The identity-theft corpus comprises deduplicated complaint narratives and includes both positive and negative labels.
Source: E018
The datasets were derived from publicly available, privacy-scrubbed consumer narratives whose submitters had consented to public release.
Source: E007
Method
Baseline
The baseline uses the same prompt structure but omits the coding guidance.
Source: E012
Environment Sample
The evaluation compares codebook and baseline prompts across language models from the Gemini and GPT families, with repeated runs and generally default settings where feasible.
Source: E015
Metrics
Prompt performance is assessed using precision and recall.
Source: E015
Findings
For impostor scams, the codebook condition achieved strong average precision and recall and improved average precision relative to the baseline.
Source: E010
For identity theft, baseline precision was weaker, while the codebook condition attained stronger average precision and recall and improved average precision over baseline.
Source: E006
Training Setup
Development of the impostor-scam codebook included an inter-annotator agreement assessment on a shared narrative set, with strong reported agreement.
Source: E017
Limitations
Generalization may be limited to cybercrime attributes or events that are amenable to qualitative analysis and codebook development.
Source: E005
The study evaluates only a single prompt template and does not establish whether more elaborate templates would improve performance.
Source: E016
Failure Modes
Practical Implications
The findings suggest that resource-constrained organizations may be able to use pretrained models, potentially including lower-cost options, without specialized model development or domain expertise.
Source: E016
Evidence and source
Show evidence locators
Evidence labels locate support in the original paper; they do not establish independent replication.
- E001 · page 4 — 5 https://www.fdic.gov/consumer-resource-center/cybersecurity: Evidence E001
- E002 · page 3 — 1 While it is possible for an incident to involve both identity theft and an impostor: Evidence E002
- E003 · page 1 — Introduction: Evidence E003
- E004 · page 1 — Introduction: Evidence E004
- E005 · page 5 — 5 https://www.fdic.gov/consumer-resource-center/cybersecurity: Evidence E005
- E006 · page 4 — 3 Both “imposter” and “impostor” are common spellings and we did not observe any: Evidence E006
- E007 · page 2 — 1 While it is possible for an incident to involve both identity theft and an impostor: Evidence E007
- E008 · page 1 — Abstract: Evidence E008
- E009 · page 3 — 1 While it is possible for an incident to involve both identity theft and an impostor: Evidence E009
- E010 · page 4 — 3 Both “imposter” and “impostor” are common spellings and we did not observe any: Evidence E010
- E011 · page 6 — 5 https://www.fdic.gov/consumer-resource-center/cybersecurity: Evidence E011
- E012 · page 3 — 1 While it is possible for an incident to involve both identity theft and an impostor: Evidence E012
- E013 · page 3 — 1 While it is possible for an incident to involve both identity theft and an impostor: Evidence E013
- E014 · page 2 — 1 While it is possible for an incident to involve both identity theft and an impostor: Evidence E014
- E015 · page 3 — 1 While it is possible for an incident to involve both identity theft and an impostor: Evidence E015
- E016 · page 6 — 5 https://www.fdic.gov/consumer-resource-center/cybersecurity: Evidence E016
- E017 · page 3 — 1 While it is possible for an incident to involve both identity theft and an impostor: Evidence E017
- E018 · page 3 — 1 While it is possible for an incident to involve both identity theft and an impostor: Evidence E018