Help Center

Sources

Sources are the feeds and sites Mediastilo monitors for material. Every pipeline run reads your active sources, scores each item against your Brand, and surfaces the best-aligned material in Signals, ready to become drafts.

Types of source

  • RSS feeds: the most reliable kind. Paste a feed URL, or paste a site URL and Mediastilo will try to discover its feed for you.
  • Scraped sites: for sites without a usable feed, Mediastilo reads the page directly using a set of selectors that tell it where the posts, titles, and links are. It works those out for you when you add the source, you don't write them. Sites that build their article list in the browser with JavaScript are supported: Mediastilo renders the page like a real browser when the plain page text has no items, so client-side-rendered listings are read correctly.
  • Summary-only sources: some major publishers (for example large international papers) publish an open RSS feed but keep the article body behind a paywall. For these, Mediastilo ingests only the publisher's own headline and standfirst from the feed — enough to follow the story as signal, but not the full article. Because the full text isn't available, these items are not drafted from (drafting needs the real body), and they're marked Summary-only in your source list.
  • Short posts count: a four-line council notice, or a post that is a poster with no text at all, is a whole post — not a failed read. Mediastilo judges an item by whether it got everything the page had, not by how long that turned out to be, so these are usable material like any other. A post whose body we only partly hold is the one held back from drafting.

Adding a source

  1. Go to Sources and paste any URL. The site's home page is fine, and so is a feed address, a news section, or a link to a single article — you don't have to know whether the site has a feed, or where its posts live.

  2. Mediastilo works out what you gave it first. A feed address is used as the feed you meant, rather than being second-guessed. A link to one article is read as a pointer to where that article is filed, so you subscribe to the section rather than to the related posts strip at the bottom of the story.

    Then it looks for an RSS feed. If the site doesn't publish one, it looks for the list of posts itself, following the site's own navigation (blog, news, updates, objave) and the usual addresses when the navigation doesn't say. Pages that build their list in the browser are rendered like a real browser before being read.

    A feed that parses isn't always a feed that's alive — plenty of sites leave one behind after a redesign and carry on posting on the page. So when a feed's newest item is more than a week old, Mediastilo reads the site as well and keeps whichever of the two is actually current. The feed wins ties and anything close: it's one cheap request the publisher maintains, while reading the page directly depends on a layout that changes.

  3. The source is named after the site, not after its feed. A feed often calls itself something like RSS 2.0, Naslovna or the title of whichever article the address happened to point at, so Mediastilo asks the site what it calls itself instead and falls back to the domain rather than to a file format.

  4. You get what we'd collect: the most recent posts we found, with dates, the page we settled on, and the first post expandable to its full text. Confirm it, or say it's not what you meant.

  5. Mediastilo also checks whether the domain is already in your sources, so you don't add duplicates.

You're only ever asked about content, never about page structure. Whatever the site is — a newsroom, a supplier's blog, a ministry's notices — if it publishes posts in a list, it's a usable source.

A source is saved only if it actually produces posts on the spot: if the page we settled on comes back empty, or without a single date to order posts by, Mediastilo says so instead of adding a source that would quietly deliver nothing.

If we can't find a list of posts, the posts may simply live elsewhere on the site, try the section page directly (for example the site's News or Blog page). You can also still set a source up by hand, giving the selectors that say where the posts, titles, links, and dates are.

If Preview reports the site is bot-protected or behind a paywall, it can't be added automatically, some sites (many international majors) actively block automated reading. When that happens, try the site's official RSS feed if it has one; a custom scraper won't get past the same wall.

The number of sources you can add depends on your plan, see Credits & Billing.

Source health

The Sources board sorts every source into one of four buckets and puts the ones needing attention first:

  • Failing: the last fetches errored, the selectors no longer match the page, or the source answered fine and returned nothing several checks in a row.
  • Stale: nothing fetched in the last 48 hours.
  • Paused: switched off deliberately. Failures don't apply while a source is paused.
  • Healthy: fetched recently, without errors.

The counts above the list are for all your sources, not just the ones on screen, and they stay put as you switch buckets — click a tile to see only that bucket. The bucket you're looking at is part of the page address, so you can reload it, bookmark it, or share it, and the coverage note on your home page links straight to the failing ones.

Mediastilo does not ask you to repair a failing source, because almost never can you. The reasons a fetch fails - a bot wall, a paywall, a rate limit, a server error - are ours to work around, and we keep retrying on our own. What the home page tells you is the consequence: how many of your sources brought in nothing this week. That is worth knowing, because a source that has delivered nothing for a month may simply not be earning its place in your list, and dropping it is a decision only you can make.

For scraped sites Mediastilo also tracks a lower-level status, shown on the source's own page:

  • OK: the source is being read correctly.
  • Broken: the selectors no longer match the page (the site likely changed its layout).
  • Unreachable: Mediastilo couldn't reach the site at all.

A source that keeps coming back empty shows a No articles badge. It's the quiet failure: the site answers normally, so nothing errors, but nothing arrives either — usually because the site changed its layout and what the source collects no longer matches. Three empty checks in a row is enough to call it failing, so a source can't go on looking healthy while delivering nothing. Open it and run a check; if the site simply hasn't published anything, the next article clears the badge by itself.

When that happens, Mediastilo doesn't stop at naming the problem. The moment a source crosses into failing, it reads the publisher again from scratch — the same way it did when you added it, starting from their domain rather than from the address that just failed. It looks at everything they offer, not only the way we happen to read them today: a site whose feed has gone quiet can move onto reading the pages directly, and a site we read with selectors can move onto a feed it has since published. If the publisher files news in more than one place, the sections we weren't reading are added alongside rather than chosen between. If it works out how to read the new layout, the source card offers the fix. Fix it applies it and re-reads the site on the spot, so you see straight away how many posts it brought in. Nothing is offered when the layout hasn't actually changed and the page is genuinely empty: there'd be nothing to repair. Not now puts the offer away, and we'll say something again if the site changes.

A failing source also shows a Blocked badge when the site stopped the last fetch, and the badge names what stopped it. Those aren't one thing, and the source page spells out which you're looking at. A Cloudflare wall, a DataDome wall, a bot challenge or a plain bot-blocked means the publisher is refusing automated readers: there is no way past that from our side, and the source will keep failing unless they offer a feed we can read or let us through. A paywall means the article body isn't ours to take, though a feed headline and standfirst still are. A server error or rate-limited is not a wall at all — the publisher's own site had a moment, or asked us to slow down; retries back off on their own and it clears when their site does. Only the first group is worth acting on, and the action is a feed from the publisher, not more retrying. A Summary-only badge means the opposite of a failure: the source is working, but by design Mediastilo keeps only its feed headline and standfirst (a paywalled publisher), so it contributes signal without a full article to draft from.

Two more badges describe a source that is working and collecting the wrong text, which is a different problem from one that collects nothing. Noisy read means the publisher's feed carries their articles cleanly and our reading of their pages is picking up the material around the article as well — their end is fine, ours is what needs fixing. Wrong text means the feed and the article page don't match at all, which points at the wrong feed or the wrong part of the page. Both are worth telling us about: the source will look healthy on every other measure while it is happening.

If a source is broken or unreachable, fix or update it so it keeps contributing material. You can also validate a source on demand to re-check it.

What a source has produced

Health tells you whether a source is working. It can't tell you whether it's worth having — a source can sync flawlessly for months and never once produce a piece you published. So every source card also carries what it actually yielded:

collected → drafted → published

  • collected: items this source brought in.
  • drafted: distinct drafts they produced. Several items about the same story share one draft and count once, so this is stories written, not articles read.
  • published: how many of those drafts you published. This is the number that answers "is this source earning its place?"

A source that has brought items in but produced no drafts is marked No drafts yet. That isn't a fault — a source can be a useful early-warning feed without ever being drafted from — but a source sitting at zero for a long time is the first candidate to drop.

The counts respect your plan's analytics window, so they cover the same period as the rest of your reporting.

Keeping sources fresh

Sources are read automatically on every pipeline run, but you can Sync a source manually to pull its latest items right away.

Scrape configuration proposals

On higher plans, when a scraped site breaks or needs setup, Mediastilo can propose a new scrape configuration, the selectors that tell it where the content lives. You can review a proposal and confirm it (to apply it) or reject it. This keeps scraped sources working as sites change, without hand-editing selectors yourself.

Scrape-config proposals are a plan-gated feature. If you don't see them, check your plan in Credits & Billing.

Removing a source

Remove a source from the Sources list at any time. It will no longer be read on future pipeline runs.

What happens next

Material from your sources flows into Signals, where Mediastilo scores it for brand alignment, and the strongest matches become drafts. See Drafts & Publishing.

See also

  • Signals: the scored feed your sources feed into.
  • Brand: what each item is scored against.
  • Analytics: see which sources contribute the most.