The sublanguage of cross-coverage.

Peter D. Stetson; Stephen B. Johnson; Matthew Scotch; George Hripcsak

The sublanguage of cross-coverage.

Peter D. Stetson, Stephen B. Johnson, Matthew Scotch, George Hripcsak

Research output: Contribution to journal › Article › peer-review

Abstract

At Columbia-Presbyterian Medical Center, free-text "Signout" notes are typed into the electronic record by clinicians for the purpose of cross-coverage. We plan to "unlock" information about adverse events contained in these notes in a subsequent project using Natural Language Processing (NLP). To better understand the requirements for parsing, Signout notes were compared to other common medical notes (ambulatory clinic notes and discharge summaries) on a series of quantitative metrics. They are shorter (mean length 59.25 words vs. 144.11 and 340.85 for ambulatory and discharge notes respectively) and use more abbreviations (26.88% vs. 20.07% and 3.57%). Despite being terser, Signout notes use less ambiguous abbreviations (8.34% vs. 9.09% and 18.02%). Differences were found using Relative Entropy and Squared Chi-square Distance in a novel fashion to compare these medical corpora. Signout notes appear to constitute a unique sublanguage of medicine. The implications for parsing free-text cross-coverage notes into coded medical data are discussed.

Original language	English (US)
Pages (from-to)	742-746
Number of pages	5
Journal	Proceedings / AMIA ... Annual Symposium. AMIA Symposium
State	Published - 2002
Externally published	Yes

ASJC Scopus subject areas

General Medicine

Cite this

@article{32cb3506630844689b283c30a8242de4,

title = "The sublanguage of cross-coverage.",

abstract = "At Columbia-Presbyterian Medical Center, free-text {"}Signout{"} notes are typed into the electronic record by clinicians for the purpose of cross-coverage. We plan to {"}unlock{"} information about adverse events contained in these notes in a subsequent project using Natural Language Processing (NLP). To better understand the requirements for parsing, Signout notes were compared to other common medical notes (ambulatory clinic notes and discharge summaries) on a series of quantitative metrics. They are shorter (mean length 59.25 words vs. 144.11 and 340.85 for ambulatory and discharge notes respectively) and use more abbreviations (26.88% vs. 20.07% and 3.57%). Despite being terser, Signout notes use less ambiguous abbreviations (8.34% vs. 9.09% and 18.02%). Differences were found using Relative Entropy and Squared Chi-square Distance in a novel fashion to compare these medical corpora. Signout notes appear to constitute a unique sublanguage of medicine. The implications for parsing free-text cross-coverage notes into coded medical data are discussed.",

author = "Stetson, {Peter D.} and Johnson, {Stephen B.} and Matthew Scotch and George Hripcsak",

year = "2002",

language = "English (US)",

pages = "742--746",

journal = "Proceedings / AMIA ... Annual Symposium. AMIA Symposium",

issn = "1531-605X",

publisher = "Hanley & Belfus",

}

TY - JOUR

T1 - The sublanguage of cross-coverage.

AU - Stetson, Peter D.

AU - Johnson, Stephen B.

AU - Scotch, Matthew

AU - Hripcsak, George

PY - 2002

Y1 - 2002

N2 - At Columbia-Presbyterian Medical Center, free-text "Signout" notes are typed into the electronic record by clinicians for the purpose of cross-coverage. We plan to "unlock" information about adverse events contained in these notes in a subsequent project using Natural Language Processing (NLP). To better understand the requirements for parsing, Signout notes were compared to other common medical notes (ambulatory clinic notes and discharge summaries) on a series of quantitative metrics. They are shorter (mean length 59.25 words vs. 144.11 and 340.85 for ambulatory and discharge notes respectively) and use more abbreviations (26.88% vs. 20.07% and 3.57%). Despite being terser, Signout notes use less ambiguous abbreviations (8.34% vs. 9.09% and 18.02%). Differences were found using Relative Entropy and Squared Chi-square Distance in a novel fashion to compare these medical corpora. Signout notes appear to constitute a unique sublanguage of medicine. The implications for parsing free-text cross-coverage notes into coded medical data are discussed.

AB - At Columbia-Presbyterian Medical Center, free-text "Signout" notes are typed into the electronic record by clinicians for the purpose of cross-coverage. We plan to "unlock" information about adverse events contained in these notes in a subsequent project using Natural Language Processing (NLP). To better understand the requirements for parsing, Signout notes were compared to other common medical notes (ambulatory clinic notes and discharge summaries) on a series of quantitative metrics. They are shorter (mean length 59.25 words vs. 144.11 and 340.85 for ambulatory and discharge notes respectively) and use more abbreviations (26.88% vs. 20.07% and 3.57%). Despite being terser, Signout notes use less ambiguous abbreviations (8.34% vs. 9.09% and 18.02%). Differences were found using Relative Entropy and Squared Chi-square Distance in a novel fashion to compare these medical corpora. Signout notes appear to constitute a unique sublanguage of medicine. The implications for parsing free-text cross-coverage notes into coded medical data are discussed.

UR - http://www.scopus.com/inward/record.url?scp=0036369894&partnerID=8YFLogxK

UR - http://www.scopus.com/inward/citedby.url?scp=0036369894&partnerID=8YFLogxK

M3 - Article

C2 - 12463923

AN - SCOPUS:0036369894

SN - 1531-605X

SP - 742

EP - 746

JO - Proceedings / AMIA ... Annual Symposium. AMIA Symposium

JF - Proceedings / AMIA ... Annual Symposium. AMIA Symposium

ER -

The sublanguage of cross-coverage.

Abstract

ASJC Scopus subject areas

Other files and links

Fingerprint

Cite this