We're sorry but this app doesn't work properly without JavaScript enabled. Please enable it to continue.

This lesson's interactive features are locked, please to keep using them

Extract Page Content

Our web crawler will need to know how to read a page of HTML. Now that we can normalize URLs, let's start extracting actual content from web pages. A web scraper that only normalizes URLs isn't very useful - we need to parse HTML and extract the meaningful information.

Let's start with just parsing the <h1> and first <p> tags.

For example, from this HTML page:

<html>
  <body>
    <h1>Welcome to Boot.dev</h1>
    <main>
      <p>Learn to code by building real projects.</p>
      <p>This is the second paragraph.</p>
    </main>
  </body>
</html>

We want to extract:

  • <h1>: "Welcome to Boot.dev"
  • The first <p>: "Learn to code by building real projects."

Assignment

func getHeadingFromHTML(html string) string
  • html is an HTML string
  • In your tests make sure that the function returns the text content of the <h1> tag if present, or the <h2> tag as a fallback.
  • Returns an empty string if neither an <h1> nor an <h2> tag is found.

Install the dependency first:

go get github.com/PuerkitoBio/goquery

I'll try not to give too many hints: read the package docs for references. Reading docs is vital practice, that said, here are a few hints:

  • .Find() scans the document looking for the passed in tag. Check the returned type.
  • The goquery.NewDocumentFromReader expects an io.Reader type to read data from. You can use strings.NewReader(html) to satisfy that.
func getFirstParagraphFromHTML(html string) string
  • In your tests make sure that the tests return the text content of the first <p> tag.
  • Returns an empty string if no <p> tag is found.

You may find that the first <p> tag doesn't always have the best or most useful results. I'd recommend searching for the <main> tag if it exists and find the first <p> tag within it, if it doesn't exist fallback to just
the first <p> tag.

Here are some example test cases to get you started:

func TestGetHeadingFromHTMLBasic(t *testing.T) {
	inputBody := "<html><body><h1>Test Title</h1></body></html>"
	actual := getHeadingFromHTML(inputBody)
	expected := "Test Title"

	if actual != expected {
		t.Errorf("expected %q, got %q", expected, actual)
	}
}

func TestGetFirstParagraphFromHTMLMainPriority(t *testing.T) {
	inputBody := `<html><body>
		<p>Outside paragraph.</p>
		<main>
			<p>Main paragraph.</p>
		</main>
	</body></html>`
	actual := getFirstParagraphFromHTML(inputBody)
	expected := "Main paragraph."

	if actual != expected {
		t.Errorf("expected %q, got %q", expected, actual)
	}
}

Run and submit the CLI tests from the root of your module.