Percorrer por autor "Macedo, Joaquim"
A mostrar 1 - 7 de 7
Resultados por página
Opções de ordenação
- It Is the Time for Portuguese Texts!Publication . Craveiro, Olga; Macedo, Joaquim; Madeira, HenriqueIn this work, we introduce a software testbed for temporal processing of Portuguese texts, composed by several building blocks: identification, classification and resolution of temporal expressions and temporal text segmentation. Starting from a simple document, we can reach a set of temporally annotated segments, which enables the establishment of relationships between words and time. This temporally enriched information is then placed into an Information Retrieval system. This work represents a step forward for Portuguese language processing, with notorious lack of tools. Its main novelty is temporal segmentation of texts. Even with target application in temporal aware Information Retrieval, the described software tools can be used in other application scenarios.
- Leveraging temporal expressions for segmented-based information retrievalPublication . Craveiro, Olga; Macedo, Joaquim; Madeira, HenriqueThe extraction of temporal information from text documents is becoming increasingly important in many applications such as natural language processing, information retrieval, question answering, etc. Indeed, the temporal dimension plays a key role on most of these systems, promoting better performance. Our goal is the definition of a temporal document representation, incorporating the time dimension into information retrieval model to improve the quality of the results. Our approach is based on temporal segmentation of documents. Temporal-aware retrieval models may explore a richer temporal document representation, enabled by segmentation. To achieve this, first we must identify temporal expressions and capture, when possible, their normalized time values. Starting from our prior work on temporal expressions recognition, we present in this paper, a resolution tool that achieves promising results in a Portuguese collection. Furthermore, a temporal characterization of the used collection shows enough and suitable information for a meaningful temporal document segmentation.
- Query Expansion with Temporal Segmented TextsPublication . Craveiro, Olga; Macedo, Joaquim; Madeira, HenriqueThe use of temporal data extracted from text, to improve the effectiveness of Information Retrieval systems, has recently been the focus of important research work. Our research hypothesis is that the usage of the temporal relationship between words improves the Information Retrieval results. For this purpose, the texts are temporally segmented to establish a relationship between words and dates found in texts. This approach was applied in Query Expansion systems, using a collection with Portuguese newspaper texts. The results showed that the use of the temporality of words can enhance retrieval effectiveness. In particular for time-sensitive queries, we achieved 9.5% improvement in Precision@10. To our knowledge, this is the first work using temporal text segmentation to improve retrieval results.
- Temporal analysis of CHAVE collectionPublication . Craveiro, Olga; Macedo, Joaquim; Madeira, HenriqueThe importance of temporal information is increasing in several Information Retrieval tasks.
- Time-Aware Focused Web CrawlingPublication . Pereira, Pedro; Macedo, Joaquim; Craveiro, Olga; Madeira, HenriqueThere is a plethora of information inside the Web. Even the top commercial search engines can not download and index all the available information. So, in the recent years, there are several research works on the design and implementation of focused topic crawlers and also on geographic scope crawlers. Despite other areas of information retrieval, research on Web crawling is not using the temporal information extracted from Web pages in the used crawling criteria. Therefore, our research challenge is the use of temporal data extracted from Web pages as the main crawling criteria to satisfy a given temporal focus. The importance of the time dimension is quite amplified when combined with topic or geography, but now we want to study it isolated. The used approach is based on temporal segmentation of Web pages text. It only follows links within segments tagged with dates in the scope of restriction. A precision around 75% was achieved in preliminary experimental results.
- Use of Co-occurrences for Temporal Expressions AnnotationPublication . Craveiro, Olga; Macedo, Joaquim; Madeira, HenriqueThe annotation or extraction of temporal information from text documents is becoming increasingly important in many natural language processing applications such as text summarization, information retrieval, question answering, etc.. This paper presents an original method for easy recognition of temporal expressions in text documents. The method creates semantically classified temporal patterns, using word co-occurrences obtained from training corpora and a pre-defined seed keywords set, derived from the used language temporal references. A participation on a Portuguese named entity evaluation contest showed promising effectiveness and efficiency results. This approach can be adapted to recognize other type of expressions or languages, within other contexts, by defining the suitable word sets and training corpora.
- Words Temporality for Improving Query ExpansionPublication . Craveiro, Olga; Macedo, Joaquim; Madeira, HenriqueThere is a lot of recent work aimed at improving the effectiveness in Information Retrieval results based on temporal information extracted from texts. Some works use all dates but others use only document creation or modification timestamps. However, no previous work explicitly focuses on the use of dates within in the document content to establish temporal relationships between words in the document. This work estimates these relationships through a temporal segmentation of the texts, exploring them to expand queries. It was achieved very promising results (13% improvement in Precision@15), especially for temporal aware queries. To the best of our knowledge, this is the first work using temporal text segmentation to improve retrieval results.
