html - How can I read and parse the contents of a webpage in R

Question

Welcome To Ask or Share your Answers For Others

html - How can I read and parse the contents of a webpage in R

1 Reply

深蓝 · Answer 1 · 2021-10-23T17:47:35+0000

Not really sure how you want to process that page, because it's really messy. As we re-learned in this famous stackoverflow question, it's not a good idea to do regex on html, so you will definitely want to parse this with the XML package.

Here's an example to get you started:

require(RCurl)
require(XML)
webpage <- getURL("http://www.haaretz.com/")
webpage <- readLines(tc <- textConnection(webpage)); close(tc)
pagetree <- htmlTreeParse(webpage, error=function(...){}, useInternalNodes = TRUE)
# parse the tree by tables
x <- xpathSApply(pagetree, "//*/table", xmlValue)  
# do some clean up with regular expressions
x <- unlist(strsplit(x, "
"))
x <- gsub("","",x)
x <- sub("^[[:space:]]*(.*?)[[:space:]]*$", "\1", x, perl=TRUE)
x <- x[!(x %in% c("", "|"))]

This results in a character vector of mostly just webpage text (along with some javascript):

> head(x)
[1] "Subscribe to Print Edition"              "Fri., December 04, 2009 Kislev 17, 5770" "Israel Time:??16:48??(EST+7)"           
[4] "????Make Haaretz your homepage"          "/*check the search form*/"               "function chkSearch()"

Categories

html - How can I read and parse the contents of a webpage in R

html - How can I read and parse the contents of a webpage in R

Please log in or register to add a comment.

Please log in or register to reply this article.

1 Reply

Please log in or register to add a comment.

Just Browsing Browsing

Most popular tags