

0 / 2 embers
0 / 3000 xp
click for more info
Complete a lesson to start your streak
click for more info
Difficulty: 7
click for more info
Not enough gems
Cost: 6 gems
1: Welcome
incomplete
2: Golang Setup
incomplete
3: Normalize URLs
incomplete
4: Extract Page Content
incomplete
5: Extract Links and Images
incomplete
6: Structure Page Data
incomplete
This lesson's interactive features are locked, please to keep using them
Our web crawler will need to know how to read a page of HTML. Now that we can normalize URLs, let's start extracting actual content from web pages. A web scraper that only normalizes URLs isn't very useful - we need to parse HTML and extract the meaningful information.
Let's start with just parsing the <h1> and first <p> tags.
For example, from this HTML page:
<html>
<body>
<h1>Welcome to Boot.dev</h1>
<main>
<p>Learn to code by building real projects.</p>
<p>This is the second paragraph.</p>
</main>
</body>
</html>
We want to extract:
<h1>: "Welcome to Boot.dev"<p>: "Learn to code by building real projects."func getHeadingFromHTML(html string) string
html is an HTML string<h1> tag if present, or the <h2> tag as a fallback.<h1> nor an <h2> tag is found.Install the dependency first:
go get github.com/PuerkitoBio/goquery
I'll try not to give too many hints: read the package docs for references. Reading docs is vital practice, that said, here are a few hints:
.Find() scans the document looking for the passed in tag. Check the returned type.goquery.NewDocumentFromReader expects an io.Reader type to read data from. You can use strings.NewReader(html) to satisfy that.func getFirstParagraphFromHTML(html string) string
<p> tag.<p> tag is found.You may find that the first <p> tag doesn't always have the best or most useful results. I'd recommend searching for the <main> tag if it exists and find the first <p> tag within it, if it doesn't exist fallback to just
the first <p> tag.
Here are some example test cases to get you started:
func TestGetHeadingFromHTMLBasic(t *testing.T) {
inputBody := "<html><body><h1>Test Title</h1></body></html>"
actual := getHeadingFromHTML(inputBody)
expected := "Test Title"
if actual != expected {
t.Errorf("expected %q, got %q", expected, actual)
}
}
func TestGetFirstParagraphFromHTMLMainPriority(t *testing.T) {
inputBody := `<html><body>
<p>Outside paragraph.</p>
<main>
<p>Main paragraph.</p>
</main>
</body></html>`
actual := getFirstParagraphFromHTML(inputBody)
expected := "Main paragraph."
if actual != expected {
t.Errorf("expected %q, got %q", expected, actual)
}
}
Run and submit the CLI tests from the root of your module.