AI Engineer — agents & data extraction

I turn messy websites and other unstructured sources into data agents can answer from.

Most of what a company knows sits in web pages and documents that no model can use as they are. I build the pipelines that turn them into data an AI assistant can answer from, with every answer linked to its source.

Start here

Live demo. I crawled a real company's public website and put an assistant on top of it. Ask it anything; every answer cites the page it came from.
ask.hellon.rio (going live soon)

What I've built since 2023

Before this, ten years of web products in JavaScript, React and Node, and a physics degree.

Work with me

Open to contract or full-time work, remote, with small teams building agents or data-extraction products.
LinkedIn · GitHub

Hellon Canella
AI Engineer — agents & data extraction

Your website, turned into answers an agent can cite.

I turn messy websites and other unstructured sources into data agents can answer from. Pick a company, I crawl its public site, and you get a link to an assistant that answers with sources.

Try the live demo →
going live soon
ask.hellon.rio · assistant over a public company website
Do you ship to Portugal?
Yes. Orders to Portugal ship from the EU warehouse in 3–5 business days.
source: /shipping#europe
Who do I talk to about bulk pricing?
The sales team handles orders above 100 units; the contact form is on the wholesale page.
source: /wholesale

Ingest only what changed

An incremental crawler re-processes changed pages only, so a large site stays current for almost nothing.

Strip the template, not the content

Sibling pages share a frame; it is detected per site and removed before any model runs.

Answers you can check

Every extracted fact keeps a verified link to its passage, and every context change must win an eval.

Hellon Canella

AI Engineer — agents & data extraction
LinkedIn · GitHub

I turn messy websites and other unstructured sources into data agents can answer from.

Crawl

Incremental: only pages that changed are processed again.

Strip

Site-level template detection removes boilerplate without a model.

Extract

Facts keep a verified link to the passage they came from.

Answer

A bounded agent loop; context changes must win an eval.

The case: a real company website

The pipeline above runs end to end on a public company site. The result is a chat link where every answer cites the page it came from: ask.hellon.rio (going live soon).

Built in production since 2023 for a B2B assistant platform; before that, ten years of web products in JavaScript, React and Node, and a physics degree.