At Context.dev, we're building an API that helps you fetch company brand data, like name, address, logos, colors, and more from any domain with a single API call.

Web scraping company logos from websites at scale can be a challenge, but with the right approach you can automate it. This guide covers techniques for extracting logos using Node.js and TypeScript, including DOM traversal tricks, filtering out non-logo images, handling different image formats, dealing with dynamic content, avoiding scraper blockers, and scaling up to thousands of domains. Code examples in TypeScript are provided along the way.

A brief word on an alternative CV based approach

This blog post is going to be primarily focused on DOM-parsing techniques, notes, and learned insights. However when DOM parsing fails, you can also treat the entire page as an image and apply vision algorithms. A mixture of these approaches produces the best results based on our testing.

Identifying Logo Elements in the DOM

The first step is to locate the logo in a webpage's HTML. Most websites place their main logo in the header or navigation area, typically as an <img> tag or an SVG element. Here are common patterns to find logos in the DOM:

HTML Attributes

Look for <img>, <picture>, or <object> tags with id or class attributes containing keywords like "logo", "brand", "header-logo", "site-logo", etc.

Alt Text

Check the alt attribute of images for the word "logo" or the website's name. E.g. <img src="logo.png" alt="Acme Corp Logo">.

Container Text

Check the container's text attributes for words that denote whether the image inside is a logo or not.

Filename/Path

The image src URL often contains "logo" (e.g. /images/logo.png). This isn't foolproof, but can be a clue.

Logo Link

The logo is often wrapped in a link to the homepage. For example, <a href="/" ...><img ...></a>. Finding an <a> tag linking to the root of the site with an image inside is a strong indicator.

Header Container

Logos usually reside in the header or nav section. If the site has a <header> or <nav> element, searching within it for an <img> can narrow the scope.

TypeScript Example

import * as cheerio from 'cheerio';
import axios from 'axios';

async function findLogo(url: string): Promise<string | null> {
    try {
        const response = await axios.get(url);
        const $ = cheerio.load(response.data);

// Common logo selectors
        const logoSelectors = ['img[class*="logo"]', 'img[id*="logo"]', 'img[alt*="logo"]', 'img[alt*="brand"]', 'header img', 'nav img', 'a[href="/"] img', 'a[href="/index.html"] img'];

for (const selector of logoSelectors) {
            const img = $(selector).first();
            if (img.length > 0) {
                const src = img.attr('src');
                if (src) {
                    // Convert relative URLs to absolute
                    return new URL(src, url).href;
                }
            }
        }

return null;
    } catch (error) {
        console.error('Error finding logo:', error);
        return null;
    }
}

Filtering Out Non-Logo Images

Once you've extracted candidate images, you need to ensure they are actually the real logo and not some other graphic. Here are strategies to distinguish the true logo from generic images:

Size and Aspect Ratio

Logos are usually of moderate size – neither tiny icons nor huge full-screen images.

File Path Keywords

Look at the image filename or URL path. If it contains keywords like icon, icons, social, or names of social networks, it's likely not the main logo.

HTML Element Context

Examine the DOM context. If an image is inside a <div class="carousel"> or <section class="banner">, it's probably a content image.

Anchor Link Destination

As mentioned, the logo is commonly wrapped in a link to the homepage.

By applying these filters, you can usually zero in on the real logo.

Avoiding Social Media Icons

Almost every website also has tiny icons for social media (Facebook, Twitter, LinkedIn, etc.), and you definitely want to avoid misclassifying those as logos. Here are some checks:

Size Filter

Social icons are small (often 16x16 up to 32x32 pixels).

Filename/URL

The src of social icons might contain the name of the social network or generic terms.

Alt/Title Attributes

Often these icons have alt text like "Facebook" or a title attribute like "Follow us on Twitter".

CSS Classes

The presence of classes like "social-icons", "icon-facebook", etc.

Deduplicating Logo Variations

When scraping at scale, you might encounter multiple instances of what is essentially the same logo. Here are a few techniques to deduplicate images:

Exact Duplicate Check (Hashing)

First, use a quick hash like MD5 or SHA-256 on the image data to catch exact byte-for-byte duplicates.

Perceptual Hashing (pHash)

Images that look alike produce similar or even identical pHashes.

Handling Different Image Formats

Websites use various formats for logos:

SVG (Scalable Vector Graphics)

Many modern sites use SVG for logos because it's resolution-independent and crisp on all screens.

PNG

A very common format for logos (supports transparency, which is often used).

JPEG

Not as common for logos.

GIF

Rare for logos, except maybe an older site or an animated logo.

Handling Dynamic Content

Not all websites deliver the logo in the initial HTML. Many modern sites might render the layout via JavaScript. To handle this, use a headless browser to execute the JavaScript and then extract the logo.

Avoiding Scraper Blockers

When scraping many websites, you'll inevitably run into anti-scraping measures. Consider these best practices:

Rotating Proxies/IPs

Use a pool of proxy servers or IP addresses and rotate them between requests.

User-Agent Rotation

Vary the User-Agent header in your requests.

Headless Evasion

By default, headless Chrome has indicators that some anti-bot scripts detect.

CAPTCHA and Cloudflare Bypass

In rare cases, a site might present a CAPTCHA or a JavaScript challenge before letting you in.

Scaling Up to Thousands of Domains

Scraping one site is easy; scraping thousands requires careful planning for efficiency and stability. Considerations include parallelism, using tools like Puppeteer Cluster, handling resource management, implementing timeouts, and planning for data storage.