

0 / 2 embers
0 / 3000 xp
click for more info
Complete a lesson to start your streak
click for more info
Difficulty: 4
click for more info
Not enough gems
Cost: 6 gems
1: Welcome
incomplete
2: TypeScript Setup
incomplete
3: Normalize URLs
incomplete
4: Extract Page Content
incomplete
5: Extract Links and Images
incomplete
6: Structure Page Data
incomplete
This lesson's interactive features are locked, please to keep using them
So far we've built functions to help us normalize URLs and extract links/text from HTML. Now let's structure that data in a way that's much more usable.
function extractPageData(html: string, pageURL: string): ExtractedPageData;
html is an HTML stringpageURL is the absolute URL of the page (used for converting relative URLs)url, heading, firstParagraph, outgoingLinks, imageURLs. I have created the object ExtractedPageData to return the data.Here's one example test case to get you started:
test("extractPageData basic", () => {
const inputURL = "https://crawler-test.com";
const inputBody = `
<html><body>
<h1>Test Title</h1>
<p>This is the first paragraph.</p>
<a href="/link1">Link 1</a>
<img src="/image1.jpg" alt="Image 1">
</body></html>
`;
const actual = extractPageData(inputBody, inputURL);
const expected = {
url: "https://crawler-test.com",
heading: "Test Title",
first_paragraph: "This is the first paragraph.",
outgoing_links: ["https://crawler-test.com/link1"],
image_urls: ["https://crawler-test.com/image1.jpg"],
};
expect(actual).toEqual(expected);
});
Run and submit the CLI tests to verify your extraction logic works correctly!