Admin guide
This guide is for Skylogs administrators: users and RBAC, teams, notification channels, routing, status pages, and operational maintenance. For installation see Installation; for HA and multi-zone topology see Deployment.
First steps after installation
- Change the default admin password and create your own admin account.
- Configure at least one notification channel (below) and verify it works.
- Create teams and invite users.
- Create an integration (Integrations) and fire a test alert end to end.
Users, roles, and RBAC
Skylogs implements role-based access control aligned with the shared-responsibility model: teams own their alerts and policies; global admin rights stay rare.
| Role | Scope | Typical use |
|---|---|---|
| Admin | Instance-wide | Platform owners: channels, integrations, global settings |
| Team manager | One or more teams | Owns schedules, escalation policies, team membership |
| Responder | Team | Handles alerts and incidents, edits own endpoints |
| Viewer | Team or instance | Read-only: dashboards, reports, status |
Principles worth enforcing:
- Give teams manager rights over their own space instead of adding admins.
- Use service accounts with scoped tokens for API integrations — never a human's token in automation.
- Review the audit log periodically for permission changes.
Teams
Teams are the routing and ownership unit. For each team, configure members and roles, on-call schedules (rotations, shift lengths, timezone, override approval ), and escalation policies. Every policy should end in a step that cannot be missed.
Notification channels
Configure instance-wide channels; users then attach personal endpoints on each.
- Email — SMTP settings
- SMS / Voice call — gateway configuration
- Slack — bot token, per-team channels; supports acknowledge-from-Slack
- Microsoft Teams — connector/webhook configuration
- Telegram — bot token
Endpoint verification
- On creation, the user must confirm a verification message.
- Periodic re-verification can be enabled per channel type.
- Unverified or failing endpoints are flagged on the user profile and in the team readiness view; escalation policies can skip unverified endpoints and alert the team manager.
Alert routing, deduplication, and noise control
- Routing rules map incoming alerts to teams by source, tags, severity, and resource.
- Deduplication collapses repeats by fingerprint /
dedup_key. - Correlation groups related alerts to keep responders out of alert storms.
- Maintenance windows suppress expected noise during planned work — scope them to tags/resources, never instance-wide.
Review the unrouted alerts view regularly: anything landing there has no owning team — a gap in your shared-responsibility map.
Status pages
Create public or private status pages per product/service: components, current state, incident updates, subscriber notifications. Branding (logo, domain) is configurable.
Reporting & SLA configuration
Define services and their SLA targets; Skylogs computes availability, MTTA, and MTTR per service over long periods. Schedule recurring reports to stakeholders.
Audit log
All administrative and response actions are recorded: who changed a policy, who acknowledged what, which notifications were sent where. Use it for compliance evidence and post-incident review.
Operational maintenance
- Backups — back up MongoDB data and configuration per zone; test restores. In multi-zone deployments alert data is zone-local: back up each zone.
- Upgrades — follow release notes; in HA clusters upgrade nodes one at a time. A brief Raft leader election during a rolling upgrade is normal and loses no alerts.
- Monitoring Skylogs itself — expose health endpoints to an external checker and configure a dead-man's-switch (a heartbeat alert that fires when Skylogs stops reporting). The incident platform must not be its own only observer.
- Retention — configure alert and incident retention to match compliance requirements.
Security checklist
- Default credentials changed; admin accounts minimal
- TLS on the UI/API and between zones
- API tokens scoped and rotated; service accounts for automation
- RBAC reviewed; team managers own team config
- Audit log retention configured
- Backups encrypted and restore-tested