At Context.dev, we're building an API that helps you fetch company brand data, like name, address, logos, colors, and more from any domain with a single API call.
Web scraping company logos from websites at scale can be a challenge, but with the right approach you can automate it. This guide covers techniques for extracting logos using Node.js and TypeScript, including DOM traversal tricks, filtering out non-logo images, handling different image formats, dealing with dynamic content, avoiding scraper blockers, and scaling up to thousands of domains. Code examples in TypeScript are provided along the way.
A brief word on an alternative CV based approach
This blog post is going to be primarily focused on DOM-parsing techniques, notes, and learned insights. However when DOM parsing fails, you can also treat the entire page as an image and apply vision algorithms. A mixture of these approaches produces the best results based on our testing.
Identifying Logo Elements in the DOM
The first step is to locate the logo in a webpage's HTML. Most websites place their main logo in the header or navigation area, typically as an <img> tag or an SVG element. Here are common patterns to find logos in the DOM:
HTML Attributes
Look for <img>, <picture>, or <object> tags with id or class attributes containing keywords like "logo", "brand", "header-logo", "site-logo", etc.
Alt Text
Check the alt attribute of images for the word "logo" or the website's name. E.g. <img src="logo.png" alt="Acme Corp Logo">.
Container Text
Check the container's text attributes for words that denote whether the image inside is a logo or not.
Filename/Path
The image src URL often contains "logo" (e.g. /images/logo.png). This isn't foolproof, but can be a clue.
Logo Link
The logo is often wrapped in a link to the homepage. For example, <a href="/" ...><img ...></a>. Finding an <a> tag linking to the root of the site with an image inside is a strong indicator.
Header Container
Logos usually reside in the header or nav section. If the site has a <header> or <nav> element, searching within it for an <img> can narrow the scope.
TypeScript Example
import * as cheerio from 'cheerio';
import axios from 'axios';
async function findLogo(url: string): Promise<string | null> {
try {
const response = await axios.get(url);
const $ = cheerio.load(response.data);
// Common logo selectors
const logoSelectors = ['img[class*="logo"]', 'img[id*="logo"]', 'img[alt*="logo"]', 'img[alt*="brand"]', 'header img', 'nav img', 'a[href="/"] img', 'a[href="/index.html"] img'];
for (const selector of logoSelectors) {
const img = $(selector).first();
if (img.length > 0) {
const src = img.attr('src');
if (src) {
// Convert relative URLs to absolute
return new URL(src, url).href;
}
}
}
return null;
} catch (error) {
console.error('Error finding logo:', error);
return null;
}
}
Filtering Out Non-Logo Images
Once you've extracted candidate images, you need to ensure they are actually the real logo and not some other graphic. Here are strategies to distinguish the true logo from generic images:
Size and Aspect Ratio
Logos are usually of moderate size – neither tiny icons nor huge full-screen images.
File Path Keywords
Look at the image filename or URL path. If it contains keywords like icon, icons, social, or names of social networks, it's likely not the main logo.
HTML Element Context
Examine the DOM context. If an image is inside a <div class="carousel"> or <section class="banner">, it's probably a content image.
Anchor Link Destination
As mentioned, the logo is commonly wrapped in a link to the homepage.
By applying these filters, you can usually zero in on the real logo.
Avoiding Social Media Icons
Almost every website also has tiny icons for social media (Facebook, Twitter, LinkedIn, etc.), and you definitely want to avoid misclassifying those as logos. Here are some checks:
Size Filter
Social icons are small (often 16x16 up to 32x32 pixels).
Filename/URL
The src of social icons might contain the name of the social network or generic terms.
Alt/Title Attributes
Often these icons have alt text like "Facebook" or a title attribute like "Follow us on Twitter".
CSS Classes
The presence of classes like "social-icons", "icon-facebook", etc.
Deduplicating Logo Variations
When scraping at scale, you might encounter multiple instances of what is essentially the same logo. Here are a few techniques to deduplicate images:
Exact Duplicate Check (Hashing)
First, use a quick hash like MD5 or SHA-256 on the image data to catch exact byte-for-byte duplicates.
Perceptual Hashing (pHash)
Images that look alike produce similar or even identical pHashes.
Handling Different Image Formats
Websites use various formats for logos:
SVG (Scalable Vector Graphics)
Many modern sites use SVG for logos because it's resolution-independent and crisp on all screens.
PNG
A very common format for logos (supports transparency, which is often used).
JPEG
Not as common for logos.
GIF
Rare for logos, except maybe an older site or an animated logo.
Handling Dynamic Content
Not all websites deliver the logo in the initial HTML. Many modern sites might render the layout via JavaScript. To handle this, use a headless browser to execute the JavaScript and then extract the logo.
Avoiding Scraper Blockers
When scraping many websites, you'll inevitably run into anti-scraping measures. Consider these best practices:
Rotating Proxies/IPs
Use a pool of proxy servers or IP addresses and rotate them between requests.
User-Agent Rotation
Vary the User-Agent header in your requests.
Headless Evasion
By default, headless Chrome has indicators that some anti-bot scripts detect.
CAPTCHA and Cloudflare Bypass
In rare cases, a site might present a CAPTCHA or a JavaScript challenge before letting you in.
Scaling Up to Thousands of Domains
Scraping one site is easy; scraping thousands requires careful planning for efficiency and stability. Considerations include parallelism, using tools like Puppeteer Cluster, handling resource management, implementing timeouts, and planning for data storage.