October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsPC HealthRecommendedCrashes, freezes, slowdowns? Check your PC nowSpot repairable issues before they interrupt work.Check PCOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
The Finance Base
The Money Desk · Blog
Re:

How S&P Global Uses Web Crawling, Ensemble Learning and Snowflake to Expand SME Risk Coverage Fivefold

S&P Global’s reported RiskGauge architecture combines multi-layer website crawling, text extraction, ensemble learning and Snowflake processing to expand U.S. private-SME coverage fivefold, while leaving important questions about data quality, governance and predictive performance.
From TheFinanceBase Team7 min to read
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

S&P Global says it expanded RiskGauge coverage of U.S. private small and midsize enterprises (SMEs) from about 2 million to roughly 10 million—five times as many companies—by combining multi-layer website crawling, text extraction, ensemble machine-learning models, anonymized third-party data and Snowflake-based processing. The increase is a coverage claim, not evidence of a fivefold improvement in predictive accuracy.

The system turns public web signals into firmographic inputs for risk scores. It can help lenders, insurers, procurement teams and investors assess companies with limited conventional disclosure, but it does not replace financial statements, payment records or human review.

Why private SMEs are difficult to assess

Public companies usually provide standardized financial disclosures. Many private SMEs do not publish quarterly accounts or investor-grade reporting, so a lender or large buyer may lack basic evidence about a company’s existence, activity, sector, location, scale and momentum.

The problem is often coverage rather than total invisibility. A legitimate business may appear in registries, local news, social profiles, filings or a website, but those sources are fragmented and inconsistent. S&P told VentureBeat that its target was approximately 10 million active U.S. private SMEs, excluding sole proprietorships, compared with about 2 million previously covered. Those figures are reported estimates, not an independently audited count of every U.S. SME.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What the fivefold expansion means

RiskGauge is designed to combine financial, business and market-risk information into scores, reports, monitoring and comparisons. S&P describes a 1-to-100 scale in which 1 is the highest credit risk and 100 the lowest. Its current RiskGauge Desktop materials position the product for customer, supplier and counterparty analysis.

The reported fivefold result primarily means more companies receive a usable record and score. It does not establish five times more data per company, five times better default prediction or equal data depth for all 10 million businesses.

“Deep web” in this implementation

Here, “deep” mainly describes crawling several layers inside a company’s domain rather than stopping at the home page. The interview refers to landing pages, contact pages, company descriptions, news and announcement pages and other text-bearing pages, along with anonymized third-party datasets.

That is different from claiming access to password-protected, paywalled or dark-web sources. The public description does not establish that restricted databases were accessed.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The pipeline from websites to risk scores

Company domains
      ↓
Crawlers and scrapers
      ↓
Preprocessing and text extraction
      ↓
Data-mining algorithms
      ↓
Curation and validation
      ↓
Firmographic drivers
      ↓
RiskGauge scoring
      ↓
Reports, monitoring, APIs and Snowflake delivery

1. Crawling multiple URL layers

Crawlers discover and fetch pages within company domains. The approach avoids relying entirely on rigid sitemaps or site-specific rules because websites vary widely. S&P discussed processing information from more than 200 million websites or pages and several terabytes of website data; the interview uses both terms, so the precise unit is not clear.

Rank #2
Sale
Storytelling with Data: A Data Visualization Guide for Business Professionals
  • Wiley
  • Language: english
  • Book - storytelling with data: a data visualization guide for business professionals

2. Preprocessing the content

The reported preprocessing layer removes material the models do not need, particularly HTML tags, JavaScript, TypeScript and other code, while retaining human-readable text. Normalizing pages this way makes it easier to process many unrelated site designs.

This does not prove that every structured element, image, PDF or table is discarded. The public account does not provide a complete retained-data schema.

3. Mining and curating signals

Specialized algorithms extract attributes such as company name, business description, sector, location and evidence of operating activity. Curators and validation logic then prepare those outputs as inputs to RiskGauge.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

4. Scoring the company

The web-derived variables become one evidence layer alongside financial risk, business risk, market risk, business-credit information, historical performance, key developments and peer comparisons. A website is not itself a credit score.

Where ensemble learning fits

S&P reported using multiple algorithms that examine different parts of a company’s pages and “vote” on conclusions. The models reportedly help identify company attributes and detect sentiment or polarity in announcements.

Ensembling can make a system less dependent on one classifier’s weaknesses. Agreement among models may also act as a quality signal when wording and page layouts differ. However, S&P has not publicly disclosed the model families, weights, thresholds, calibration method or evaluation statistics. The description does not establish whether the implementation uses random forests, boosting, stacking, neural models or a hybrid.

Snowflake’s role

According to the interview, Snowflake’s warehouse and Snowpark Container Services sit in the preprocessing, mining and curation stages. That provides a scalable environment for large text volumes, repeated transformations, custom algorithms and downstream integration.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Snowflake is infrastructure, not the proprietary RiskGauge methodology. The fivefold coverage result reflects the combined system: discovery, extraction, model design, third-party data, operations and product integration. Snowflake offers consumption-based pricing across Standard, Enterprise and Business Critical editions; costs depend on cloud, region, edition and workload. See the pricing page and service-consumption table.

How records stay current

S&P described weekly scans, with downstream updates triggered by detected changes rather than rewriting every record every week:

  1. A hash is created for a previously captured landing page.
  2. A later crawl creates a new hash.
  3. A matching hash indicates no detected content change.
  4. A mismatch triggers further processing or an update.

Hashing controls compute and reduces unnecessary model runs, but it detects content change rather than truth. A redesign can create a change event without a meaningful business development; a dynamic page can create repeated false positives; and an unchanged page can coexist with deteriorating finances. Website-update frequency is therefore a heuristic for activity, not proof of solvency.

The accuracy-versus-speed trade-off

S&P said some algorithms had strong accuracy, precision and recall but were too expensive to run at the required scale. Production choices therefore balance:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Precision against recall and false-positive risk.
  • Model sophistication against latency and compute cost.
  • Crawl depth against coverage and expense.
  • Freshness against the number of rescoring jobs.
  • Generalization against the diversity of websites and sectors.

The practical lesson is that the best laboratory model may not be the best model for hundreds of millions of pages. Avoiding hard-coded rules and robotic process automation can improve resilience, but it also makes entity resolution and quality control harder.

What web signals can—and cannot—tell a lender

Signal type Useful indication What it does not prove
Company name, location and description Identity and basic firmographics Legal ownership or financial strength
Service and sector language Business activity and industry classification Revenue, margins or market share
Announcements and sentiment Potential developments or momentum Liquidity or repayment capacity
Recent website updates A possible activity indicator Continued operation or solvency
Rich online presence More extractable evidence Lower default probability

Web data can expand the evidence available for thin-file companies. It cannot substitute for cash flow, debt, liens, payment history, audited accounts or verified ownership. Marketing claims may be inflated, copied or maintained by an agency after a business has stopped trading.

Failure modes buyers should test

  • Stale or thin websites: legitimate firms may have little online presence, while closed firms may leave sites online.
  • Entity confusion: subsidiaries, franchises, trade names and similarly named companies can be conflated.
  • Template contamination: website-builder boilerplate can make unrelated companies look alike.
  • Dynamic or blocked pages: client-rendered content, rate limits and bot defenses can reduce coverage.
  • Sector and language bias: digital-first or English-language businesses may generate richer signals.
  • Proxy bias: web activity can reflect marketing budget, size or digitization rather than creditworthiness.
  • Concept drift: language, designs and business practices change over time.
  • Adversarial content: companies could optimize pages to influence automated classifiers.

Ensemble agreement does not remove shared bias: several models can reach the same wrong conclusion from the same misleading page.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Governance and responsible use

The public account does not answer several questions that a regulated or high-stakes deployment must address:

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Whether crawlers follow robots.txt and site terms in every jurisdiction.
  • How personal information and employee data are filtered.
  • How source URLs, timestamps, extraction methods, confidence and model versions are retained.
  • How duplicate entities and parent-subsidiary relationships are resolved.
  • How a company can correct an error or challenge an assessment.
  • Whether scores support analysts or trigger automated decisions.
  • How legal events, sanctions, adverse news and politically exposed persons are handled.

S&P reported no human in the initial ensemble decision process. That does not establish a human-free customer workflow; exception handling, complaints, model governance and material credit decisions may still require people.

Buy RiskGauge, use an API or build the stack?

Option Best fit Main trade-off
RiskGauge Desktop Enterprises managing large customer, supplier or counterparty portfolios Packaged scores and monitoring require a demo and commercial agreement; no public self-serve price is shown on the reviewed page.
Universal Coverage API Teams embedding entity matching, coverage or RiskGauge reports in underwriting or procurement software Integration control, but account-based commercial engagement; no public price was displayed on the reviewed Marketplace page.
Snowflake-based build Organizations with data-engineering capacity and proprietary data to combine Elastic processing and governance, but consumption costs and operating complexity.
Complete custom crawler and model stack Very large-volume users with a defensible need for proprietary features Maximum control requires continual work on compliance, entity resolution, lineage, monitoring, quality and dispute processes.

A custom build also needs crawl politeness, rate limiting, browser rendering where necessary, source retention, model monitoring and cost controls. Buying finished intelligence is usually simpler when the goal is immediate portfolio decisions; building makes more sense when proprietary data, workflow integration or unusual coverage requirements justify the ongoing expense.

What remains undisclosed

The available public material does not provide AUC, KS, Brier scores, calibration statistics, default-prediction lift, false-positive rates, out-of-time validation, sector-by-sector results or a comparison of web-only with web-plus-financial models. It therefore supports a description of architecture and coverage expansion, not a claim of predictive superiority.

For buyers, the most important diligence requests are model-performance evidence, coverage definitions, provenance and timestamps, correction procedures, score explanations, refresh behavior and the role of human analysts in consequential decisions.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The Bottom Line

S&P’s project shows how industrial-scale crawling and machine learning can bring millions of thin-file SMEs into a risk workflow. Its durable lesson is architectural: public web signals can broaden coverage when they are extracted, validated and monitored at scale. They remain proxies, however, and should complement—not replace—verified financial and payment evidence.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

More post from the Money Desk

  1. The Money DeskBlogTheFinanceBase07 MAR 2625 minWhat Is a 457 Plan?
  2. The Money DeskBlogTheFinanceBase07 MAR 2621 minTime Value of Money: What It Is and How It Works
  3. The Money DeskBlogTheFinanceBase07 MAR 2627 minAre You Living in One of These Top 10 Most Expensive Cities to Retire?
Recommended PC Tool
Recommended PC Tool
Windows Errors? Fix Them Before They SpreadFree repair scan
Crashes, No Sound, or Screen Glitches?Free driver scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.