Repository logo
 

Protecting Privacy in the Archives: Preliminary Explorations of Topic Modeling for Born-Digital Collections

Date

2017-12

Authors

Hutchinson, Tim

Journal Title

Journal ISSN

Volume Title

Publisher

IEEE

ORCID

Type

Conference Presentation
Refereed Paper

Degree Level

Abstract

Natural language processing (NLP) is an area of increased interest for digital archivists, although most research to date has focused on digitized rather than born-digital collections. This study in progress explores whether NLP techniques can be used effectively to surface documents requiring restrictions due to their personal information content. This phase of the research focuses on using topic modeling to find records relating to human resources. Early results show some promise, but suggest that topic modeling on its own will not be sufficient; other techniques to be explored include sentiment analysis and named entity extraction.

Description

Keywords

Citation

Tim Hutchinson, 2017. Protecting Privacy in the Archives: Preliminary Explorations of Topic Modeling for Born-Digital Collections. Proceedings of the 2017 IEEE International Conference on Big Data. Boston, MA: 11-14 December 2017, pp. 2251-2255.

Degree

Department

Program

Advisor

Committee

Citation

Tim Hutchinson, 2017. Protecting Privacy in the Archives: Preliminary Explorations of Topic Modeling for Born-Digital Collections. Proceedings of the 2017 IEEE International Conference on Big Data. Boston, MA: 11-14 December 2017, pp. 2251-2255.

Part Of

item.page.relation.ispartofseries

DOI

item.page.identifier.pmid

item.page.identifier.pmcid