Logo do repositório
 
A carregar...
Miniatura
Publicação

Time-Aware Focused Web Crawling

Utilize este identificador para referenciar este registo.
Nome:Descrição:Tamanho:Formato: 
Preview_Time-Aware Focused Web Crawling.pdf140.11 KBAdobe PDF Ver/Abrir

Orientador(es)

Resumo(s)

There is a plethora of information inside the Web. Even the top commercial search engines can not download and index all the available information. So, in the recent years, there are several research works on the design and implementation of focused topic crawlers and also on geographic scope crawlers. Despite other areas of information retrieval, research on Web crawling is not using the temporal information extracted from Web pages in the used crawling criteria. Therefore, our research challenge is the use of temporal data extracted from Web pages as the main crawling criteria to satisfy a given temporal focus. The importance of the time dimension is quite amplified when combined with topic or geography, but now we want to study it isolated. The used approach is based on temporal segmentation of Web pages text. It only follows links within segments tagged with dates in the scope of restriction. A precision around 75% was achieved in preliminary experimental results.

Descrição

Part of the book series: Lecture Notes in Computer Science (LNISA,volume 8416).

Palavras-chave

Web crawling temporal text segmentation temporal information extraction temporal information retrieval

Contexto Educativo

Citação

Pereira, P., Macedo, J., Craveiro, O., Madeira, H. (2014). Time-Aware Focused Web Crawling. In: de Rijke, M., et al. Advances in Information Retrieval. ECIR 2014. Lecture Notes in Computer Science, vol 8416. Springer, Cham. https://doi.org/10.1007/978-3-319-06028-6_53

Projetos de investigação

Unidades organizacionais

Fascículo

Editora

Springer Nature

Licença CC

Sem licença CC

Métricas Alternativas