Semeris Rockies
Semeris uses a crawler called SemerisRockies to collect documents for structured finance and credit agreement deals. If you see it in your logs, the requests came from us. The people who run it read crawler@semeris.com.
It collects indentures, offering circulars, notices to noteholders, and trustee reports for structured finance and credit agreements. Its sources are stock exchanges that list structured finance and credit agreement notes, and the trustee and administrator portals where Semeris has a login. It identifies itself on every request and generates about as much traffic as one attentive person.
What it collects, and what it doesn't
SemerisRockies fetches structured finance and credit agreement deal documents, and nothing its login is not already entitled to see.
Public exchange sites. It reads the announcement and document listings for structured finance and credit agreement issuers. Then it downloads the notices and offering documents those listings link to.
Portals that need a login. It signs in with an account issued to Semeris and lists the deals that the account can see. Then it downloads the documents and reports posted to those deals.
It does not:
try to get past an access control. If the account can't open something, it stops there;
connect to any host that isn't on the fixed list we keep for your site, or over anything but HTTPS;
pretend to be a web browser, rotate IP addresses, route through proxy networks, or solve CAPTCHA.
How to recognize it
Every request sends this user agent, with the crawler's version number in place of <version>:
Requests come from a small, fixed set of IP addresses. To check a request or add us to an allowlist, email us and we'll send you the list. If a request uses this name but comes from another address, it is not from us. Please tell us about it.
How it behaves on your site
SemerisRockies runs once a day and sends one request at a time, with a several-second wait between requests.
Signing in. If sign-in is refused several times in a row, it stops using that account until someone at Semeris has looked into it. It will not lock an account by retrying. If a portal reports that the password has expired, it stops straight away.
Errors. When your server returns a 429 or a 5xx error, the crawler does not retry that request in the same run.
robots.txt. SemerisRockies doesn't explore sites, so it doesn't consult robots.txt. For each site, it visits a set of listing pages we configure by hand, plus the documents they link to. The quickest way to slow it down or stop it is to email us.
SemerisRockies/<version> (+https://semeris.com/crawler; crawler@semeris.com)| Portals that need a login | Public exchange sites | |
|---|---|---|
| How often | Once a day | Once a day |
| Requests in flight | 1 | 1 |
| Wait between requests | 4–11 seconds, varied at random | 1.5–4 seconds, varied at random |
| Longer pauses | 2–5 minutes, about every 40 requests | 2–5 minutes, about every 40 requests |
| Most requests in any hour | 600, or fewer where a site needs it | 1,800 |
Contact us
Email crawler@semeris.com about anything SemerisRockies does on your site.
Please tell us which site it is, roughly when you saw the traffic, and what you'd like. That might be a slower pace, a different time of day, a dedicated account, or for us to stop. We'll pause crawling of your site while we sort it out.
Do you offer email delivery, a bulk download or a scheduled transfer of the same documents? If so, we'd rather use it than crawl.