Monitoring and audit
Know it stopped
before the business does.
The front that looks after what is already in production: it detects the outage, explains the failure and keeps the history of what changed in the tenant.
Not just a tool
Configuring the critical interfaces and setting the silence windows is done together with your team, in the first week. Badly calibrated monitoring means too many alerts — and too many alerts means a muted channel, which is worse than no alerting at all.
What is inside
Each capability, what it changes and why it is different
Heartbeat
Tracks every critical interface and fires when it passes the configured silence window — 5, 30 or 120 minutes, set per interface.
Gain · the time between the outage and the alert stops depending on someone opening a dashboard.
What makes it different · it detects absence of success, not presence of error. That is the only way to see the outage that raises no error — and most of them raise none.
Alerts in your team's channel
Slack, Teams, email, Telegram or WhatsApp, in each person's language, with a test button that sends a real message before you depend on it.
Gain · the warning arrives where the team already is, at 3am, with nobody watching a screen.
What makes it different · alerts are split by audience: an outage wakes the on-call rota, billing and quota go by email only. Mixing them is the shortest path to a muted channel.
Root cause (RCA)
From an outage event, it gathers the interface history, the artefact state, the latest executions and recent changes, and writes the diagnosis as a PDF.
Gain · the diagnosis stops depending on the person with five years in the company.
What makes it different · it comes out as a document ready to circulate — including with people outside integration, who are usually the ones asking for the explanation.
Operations monitor
The tenant's processing on one screen: completed, failed, processing and retrying, with duration and status per message.
Gain · answers “is it running?” without opening Cloud Integration.
What makes it different · it puts executions, certificates and queues in one place, which in CPI are three different screens.
Error Monitor
Tracks tenant failures and groups the recurring ones instead of listing the same thing a hundred times.
Gain · separates the new error from the usual one — which is what decides where the next hour goes.
What makes it different · suppression is counted and shown: you know how many times it repeated, without receiving a hundred alerts.
Integration Health
Scores tenant health across performance, security, error handling and readability, and shows where the risk is concentrated.
Gain · gives you a number for the governance meeting, with the technical detail behind it.
What makes it different · the score comes from reading the real artefact, not from a questionnaire answered by whoever built it.
Tenant audit
Keeps a snapshot of your artefacts and shows the difference between two versions: what changed, when and in which environment.
Gain · answers “what changed since the last time it worked?” without relying on anyone's memory.
What makes it different · it compares versions with a written explanation, not just an XML diff nobody reads.
Certificates
Tracks tenant certificate expiry and warns ahead of time.
Gain · removes a whole category of incidents that have a known date.
What makes it different · the warning goes by email, not through the on-call channel: it has a deadline, it does not need to wake anyone.
Troubleshoot and Incident Assistant
Answers “why did this message fail” from the real log, and guides the incident from symptom to hypothesis.
Gain · shortens first-line handling — and what escalates arrives with context.
What makes it different · it records what has already been ruled out, so today's on-call engineer builds on yesterday's work.
Continuous QA
Checks interfaces against quality rules continuously and warns when a change makes things worse.
Gain · quality stops being a quarterly audit and becomes something you follow.
What makes it different · the ruler is the same validator that rejects artefacts on import — what it asks here is what the tenant would ask there.
Retention you define
Each organisation chooses how long it keeps noise, errors and outages.
Gain · history lasts as long as your policy says, not as long as the vendor decided.
What makes it different · minimum floors protect the incident history: you cannot accidentally delete what the next audit would need.
See it working on your tenant
A guided 14-day evaluation, on your own interfaces — not on a prepared demo.