Ph.D. Thesis Defense Announcement: Han Wang: Knowledge Base Construction from Scientific Literature

Printer-friendly version

Ph.D. Thesis Defense Announcement: Han Wang: Knowledge Base Construction from Scientific LiteratureNovember 14, 2016
Ph.D. Thesis Defense Announcement

Knowledge Base Construction from Scientific Literature

Han Wang

Wednesday, November 16, 2016
Winslow Building Room 1140

Knowledge Bases (KBs) have become a functional utility as a repository of information for both humans and software agents to seek confirmed facts about the world. With the wide-ranging application of KBs, automatically constructing either generic KBs or domain-specific KBs using information extracted from multiple sources such as web pages, reports, and research papers has grown into an interesting task for both academia and industry.

This dissertation presents SciKB, an end-to-end Knowledge Base Construction system, which takes in a collection of research articles within a certain scientific domain and outputs a domain-specific KB. The resultant KB contains fact triples extracted from the input documents as well as hierarchical clusters of the entities and relations involved in the facts. Each cluster aggregates entities or relations with similar semantic meanings, and the hierarchies serve as an implicit schema of the KB.

SciKB adopts an open information extraction approach to extract fact triples from the input documents, then jointly learns the distributed representations of the involved entities and relations in an unsupervised fashion, and finally utilizes the obtained representations to organize the entities and relations into hierarchical clusters. Experiments are conducted to evaluate each component of the SciKB pipeline and the results demonstrate its effectiveness in two scientific domains: Biomedical Science and Earth Science.