start-ver=1.4 cd-journal=joma no-vol=109 cd-vols= no-issue=234 article-no= start-page=13 end-page=18 dt-received= dt-revised= dt-accepted= dt-pub-year=2009 dt-pub=20091009 dt-online= en-article= kn-article= en-subject= kn-subject= en-title=Construction of Annotated Corpus of Verb Meanings and Semantic Role Labels Based on Verb Thesaurus kn-title=動詞項構造シソーラスに基づく動詞語義ならびに意味役割付与データの構築 en-subtitle= kn-subtitle= en-abstract=Argument structure is widely recognized as an interface of mapping from grammatical structure of a sentence to shallow semantic structure. In English several large-scale language resources such as FrameNet, Propbank, and Dorr's LCS are proposed and each of them defines a kind of argument structure and some of them construct annotated corpora. These annotated corpora are very useful to build a statistical annotation system of semantic role labels. While in Japanese EDR provided a large-scale annotated corpus of semantic role labels; however the annotated sentences are not collected on the basis of verbs, thus it is hard to utilize the annotated corpus as a training corpus of statistical semantic role label system. Thus we propose another annotation corpus of argument structure on the basis of the Japanese Verb Thesaurus which is provided in previous work. Currently we annotated 1483 sentences for 120 verbs. In this manuscript we confirm that the problem issues of argument structure annotation, current annotation scheme, development of tool and quality of annotated corpus. kn-abstract=文法的構造を基に述語を中心として文の意味を記述する項構造が文の形式的解析から文の意味を処理するためのインターフェースとして期待されている.項構造は述語の語義と係り関係にある要素の役割で記述されるフレームであり,項構造が正しく付与できると,動詞の語義曖昧性解消ならびに同義の言い換えとどの要素が言い換え可能かまで明らかにすることが期待できる.英語ではすでに項構造を人手で付与した大規模コーパスが公開され利用されているが,日本語ではEDRで構築されたものの文書単位で付与されたため学習事例として利用が困難である.そこで本稿では語義の事例を中心に項構造を付与した意味役割付与データ(120語の動詞に対して1483文)を構築し,語義付与に起こる問題と現段階での対処について整理を行う.体系として動詞項構造シソーラスを利用した.項構造を文に付与することで体系の不備を同時に整理することを目標としている.構築した項構造タグ付きコーパスは公開する予定である. en-copyright= kn-copyright= en-aut-name=TakeuchiKoichi en-aut-sei=Takeuchi en-aut-mei=Koichi kn-aut-name=竹内孔一 kn-aut-sei=竹内 kn-aut-mei=孔一 aut-affil-num=1 ORCID= en-aut-name=MorimotoMaiko en-aut-sei=Morimoto en-aut-mei=Maiko kn-aut-name=森本真依子 kn-aut-sei=森本 kn-aut-mei=真依子 aut-affil-num=2 ORCID= affil-num=1 en-affil= kn-affil=岡山大学大学院自然科学研究科 affil-num=2 en-affil= kn-affil=岡山大学大学院自然科学研究科 END start-ver=1.4 cd-journal=joma no-vol=109 cd-vols= no-issue=234 article-no= start-page=1 end-page=5 dt-received= dt-revised= dt-accepted= dt-pub-year=2009 dt-pub=20091009 dt-online= en-article= kn-article= en-subject= kn-subject= en-title=Bio-medical Term Extraction with Morpho-Syntactic Rules on Simple Rule Language kn-title=SRLを利用した規則ベースの感染症用語抽出 en-subtitle= kn-subtitle= en-abstract=Simple rule language, rule-based term extraction, bio-medical terms, Disease surveillance system Bio-medical term extraction is a key technology for a surveillance system of epidemic disease news from the Web. In the previous work we applied statistical learning model to extract terms from the Web site. The previous approach is good at extracting terms with high precision rates; however it is weak at extracting new terms that do not exist in the training data. Since we usually have new disease names a new term extraction approach with high coverage for unknown or low-frequent terms is needed. Recently, Simple rule Language (SRL), a rule-based word extraction language, is freely available. The SRL also has an developing environment called SRL editor. Thus we are constructing rules of bio-medical terms on the several language (such as English, Japanese, Thai and Vietnam) for the multilingual disease surveillance system. In this manuscript we confirm how we construct rules to extract Japanese bio-medical terms from Japanese news articles. kn-abstract=我々は感染症情報をWeb上から集めて提示するBioCasterシステムを構築している.感染症情報は各国のローカルニュースに速報が出ることが予測されることから英語のみならず日本語を含めたアジア言語での開発を進めている.核となる技術は感染症に関する用語を記事から見つける用語抽出であるが,既存の手法では学習データを利用した統計的学習モデルを利用して構築した.しかしながら,新たな病気など学習データに無い用語が現れた際うまく獲得できないことが予測されるため規則に基づく用語抽出システムの構築を行う.規則ベースで用語を抽出するシステムとしてSRL(Simple Rule Language)が公開されており,ユーザは語構成ならびに文脈を規則で記述することで用語を抽出できる.そこで本研究では感染症情報に必要な用語についてどのようにSRL上で定義できるかについて明らかにする. en-copyright= kn-copyright= en-aut-name=ShinnouTakashi en-aut-sei=Shinnou en-aut-mei=Takashi kn-aut-name=新納貴志 kn-aut-sei=新納 kn-aut-mei=貴志 aut-affil-num=1 ORCID= en-aut-name=TakeuchiKoichi en-aut-sei=Takeuchi en-aut-mei=Koichi kn-aut-name=竹内孔一 kn-aut-sei=竹内 kn-aut-mei=孔一 aut-affil-num=2 ORCID= en-aut-name=NigelCollier en-aut-sei=Nigel en-aut-mei=Collier kn-aut-name=ナイジェルコリアー kn-aut-sei=ナイジェル kn-aut-mei=コリアー aut-affil-num=3 ORCID= affil-num=1 en-affil= kn-affil=岡山大学工学部情報工学科 affil-num=2 en-affil= kn-affil=岡山大学大学院自然科学研究科 affil-num=3 en-affil= kn-affil=国立情報学研究所 END