Improvements of a Hybrid Syllabus Search Tool by Syllabus-related Heuristics

Takayuki Sekiya, Yoshitatsu Matsuda, Kazunori Yamaguchi · 2022 IEEE Frontiers in Education Conference (FIE) · 2022

This Research Full Paper proposes a new method to collect course syllabi. A syllabus is essential information about a course in a university. Students grasp the topics covered by a course through its syllabus, and faculties understand the curriculum offered by the university by a set of syllabi. Thus, a syllabus helps to analyze educational activities. Our previous work proposed a hybrid method that combines Google API as a general keyword search engine and linear support vector machine (SVM) as content-based classification models. We could find more computer science (CS) syllabus pages than using each method alone by employing the hybrid method. This paper extends the hybrid method for finding a directory page with many links to the syllabus pages. We use the hybrid method to collect candidate directory pages and select the true directory pages from the candidates by the three heuristics: (1) Hyperlink-Induced Topic Search (HITS) score. (2) URL pattern, and (3) Content word. (1) HITS score: The relation between directory pages and syllabus pages resembles the relation between the HITS algorithm’s hubs and authorities. We use hub scores to select directory pages. (2) URL pattern: Pages with similar contents and roles share a part of their URLs. We exploit this observation to select syllabus pages. (3) Content word: We expect a syllabus page to include words of the Body of Knowledge (BOK) ‘Computing Science Curricula CS2013,’ released by the ACM and IEEE Computer Society. We use the words extracted from the CS2013 BOK to measure how each candidate syllabus page is related to CS2013. With these three heuristics, we achieved 32%, the percentage of CS syllabus pages included on candidate pages.

Read the paper · More papers on PaperTik