Command Center (NOC) Analyst
About this role
SPRINGBROOK SOFTWARE Command Center · Role Description Command Center (NOC) Analyst Command Center · Junior to mid-level · 24/7 monitoring / triage / incident communication Level Junior to mid-level Reports to Command Center Manager Direct reports None Scope First-line monitoring and triage of the Springbrook production estate — infrastructure, applications and network health across 15+ products. Hands-on balance Effectively all operations. Roughly 60% monitoring and triage, 25% incident logging and communication, 15% alert tuning and documentation. Coverage Shift-based, 24/7 rotation including nights and weekends, with a documented handover every shift. Role overview Local governments use Springbrook to bill for water, take payments at a counter, issue permits and close their books. Much of that work happens on a schedule that cannot slip. The Command Center is the function that notices when something is wrong before a customer has to tell us — and on a 24/7 estate, that is not a formality. As a Command Center Analyst you watch infrastructure, application and network health, triage what the alerts are telling you, take the first troubleshooting steps, and escalate accurately when it is beyond first line. You are also the voice of the incident while it is open: stakeholders find out what is happening because you told them. The judgement in this role is under-appreciated. Deciding that three unrelated-looking alerts are one Azure region problem, or that a nightly billing job is late rather than failed, is what separates a good analyst from a pair of eyes. We would rather you escalate a real problem early than sit on it to avoid being wrong. This is a deliberately supported level. Asking a senior engineer after twenty minutes of being stuck is good practice here, not a weakness — bring what you have already tried. The route onward is explicit: cleaner handovers, better triage, fewer false escalations, and a measurable reduction in the alert noise you own. People move from this seat into site reliability, cloud operations and security, and we would like you to. Scope of ownership The estate during your shift. If it breaks while you are watching, you are the person who noticed — or should have been. Accurate triage: what is actually affected, how badly, and who needs to know. The incident record — timeline, actions taken, current state — kept current while the incident is live, not written up afterwards. Shift handover. The next analyst starts with a complete picture or they do not; that is entirely on you. Stakeholder communication during active incidents, at the cadence the protocol requires. The signal quality of the alerts you work with. Recurring noise you tolerate is noise you own. Responsibilities Watch Monitor infrastructure, applications and network health 24/7 per shift, using established dashboards and alerting tools. Monitor Azure resource health, service alerts and Azure Monitor dashboards to detect and respond to infrastructure issues in real time. Use Azure Service Health and Resource Health tools to correlate incidents with underlying platform or region-level events. Watch the scheduled work that customers depend on — billing runs, batch jobs, integrations — and know what late looks like before it becomes failed. Triage and escalate Triage incoming alerts, perform initial troubleshooting, and escalate per defined severity and escalation procedures. Log and track incidents through resolution, ensuring accurate handoff documentation between shifts. Follow the runbook where one exists, and say clearly when the runbook is wrong rather than working around it silently. Escalate early and specifically. An unraised problem that costs a customer their billing window is a bigger problem than an escalation that turned out to be minor. Distinguish a single failing component from a platform-wide event before you page five teams. Communicate Communicate status updates to stakeholders during active incidents per established communication protocols. Write updates a non-technical reader can act on: what is affected, what we know, what we are doing, when we will next update. Keep the incident channel factual. No speculation presented as fact, no silence during a live incident. Hand over verbally as well as in writing when an incident crosses a shift boundary. Improve the signal Identify recurring alert patterns and recommend tuning to reduce noise and improve signal quality. Feed gaps back to Site Reliability Engineering — the alert that should exist, the dashboard that does not answer the question. Contribute to and correct runbooks based on what actually worked at 3 a.m. Automate or template the repetitive parts of your shift where you are able, and ask for help where you are not yet. Work AI-first Use AI assistants daily as part of normal work — this is the expected working method, not a shortcut. Lean on it where it is strongest: summarising log volume during an incident, drafting stakeholder updates and shift handovers, and explaining unfamiliar error output. Use it to help correlate — turning a burst of alerts into a candidate explanation you then verify against the actual tooling. Verify before you act or communicate. An AI-drafted incident update that is subtly wrong is worse than a slower accurate one. Watch for the specific failure modes: confident explanations of alerts it has not seen the data for, invented Azure behaviour, and plausible root causes that do not survive one look at the dashboard. Use the team's shared AI context assets, runbooks and communication templates rather than building your own private set. Never put customer or personally identifiable data, credentials, or live incident evidence into prompts to non-approved services. Technical environment Monitoring Azure Monitor, Log Analytics, Application Insights dashboards, plus third-party monitoring inherited across the acquired estate. Azure platform Azure Service Health and Resource Health, App Service, AKS, Azure Functions, Service Bus, Front Door and Application Gateway. Alerting and paging Alert rules, action groups and on-call escalation with defined severity levels and response expectations. Incident management Ticketing and incident records, documented runbooks, severity matrix, and stakeholder communication protocols. Estate .NET 10 and legacy .NET Framework services, Angular / Blazor / React front ends, PHP components, Azure SQL, SQL Server and MySQL across 15+ products. Query and scripting KQL for Log Analytics, with PowerShell or Bash for routine checks and diagnostics. Delivery context Azure DevOps CI/CD today, migrating toward GitHub and GitHub Actions, with Jira for work tracking — useful for correlating an incident with a recent deployment. AI tooling GitHub Copilot, Claude Code or equivalent used daily, with shared prompts, runbooks and context documents encoding Springbrook conventions. What we're looking for Core 1-4 years in a NOC, service desk, operations or monitoring role, ideally on a production system with real users. Comfortable working a shift rotation that includes nights and weekends, reliably. Working understanding of infrastructure and networking fundamentals — DNS, TCP/IP, load balancing, certificates — sufficient to isolate a problem. Some cloud exposure, ideally Azure. Depth is not expected; curiosity and the ability to ramp are. Ability to read a dashboard critically and query logs, or clear evidence you will learn to quickly. Clear, calm written communication under time pressure. This is a hard requirement, not a soft skill here. Disciplined documentation habits. Your handover is someone else's starting point. Bachelor's degree, relevant certification, or equivalent practical experience. Attitude and trajectory Evidence of getting better quickly — a tool, a platform or a domain you picked up faster than expected. You stay steady when several things are wrong at once, and you work the most important one first. You ask good questions: specific, after some effort, and with what you have already tried. You are honest about what you did and did not do during an incident, including mistakes. You are bothered by noise. Alerts that fire nightly and mean nothing should irritate you into fixing them. Interest in the domain. Public-sector finance software is more interesting than it sounds, and the people who care about it do better work. AI-assisted work — required Regular hands-on use of an AI assistant on real work, with an honest account of where it helped and where it misled you. The discipline to verify a generated explanation or update against the actual system before acting on it or sending it. Willingness to learn context engineering properly — this is a skill we will invest in you developing. Bonus KQL, or any real query-language experience against logs. Scripting in PowerShell, Bash or Python. Experience with a formal incident management process and severity model. Exposure to multi-tenant SaaS, ERP, payments or public-sector systems. Alert tuning or monitoring configuration you were responsible for, not just consumed. Relevant Azure certification such as AZ-900 — useful, but never a substitute for having watched something real. How we work, and who thrives here The Command Center is the front door to production, and it is treated that way — Site Reliability Engineering, Cloud Security and the delivery teams all depend on what you see and how you describe it. Shift work is demanding and we will not pretend otherwise; in exchange, the boundaries are clear, the handovers are real, and nobody expects you to carry an incident past the end of your shift alone. We do not expect you to already work at a senior level. We do expect you to own your shift, be honest about where you are, and be visibly better in six months than you are now. The traits below describe where we want you to get to — you are not expected to arrive with all of them fully formed. Ownership. You own what happens on your shift, from first alert to accurate handover. When something breaks at 2 a.m., you want to know why, and you want the root cause fixed rather than papered over. Accountability. You hold yourself to the standard before anyone else has to. You say what you will do, do it, and flag early and honestly when a situation is getting away from you. Self-starter. You find the next most valuable thing to do and start it. Blockers are problems you drive to resolution, not reasons to stop. Independent critical thinking. You question assumptions, including your own and the dashboard's. You can explain a triage decision with reasoning, push back when the evidence supports you, and change your mind when it does not. Creativity. You bring options to the table, not just problems. You look for the tuning change or runbook fix that removes a recurring alert rather than handling it again next week. What this role is not A passive screen-watching seat. Triage, correlation and communication are the substance of the job. A role where escalating is a failure. Escalating well and early is a core skill. A day-shift role. This is a 24/7 rotation, including nights and weekends. A seat where documentation is optional. Incomplete handover is the one thing that reliably makes an incident worse. A dead end. This is a route into site reliability, cloud operations and security for people who take it seriously.
Key Responsibilities
- Triage system alerts
- Perform initial troubleshooting
- Escalate critical issues
- Tune monitoring alerts
Requirements
Must have
- 1–4 years of experience
- Questioning Skills
- Overall Professionalism and Ethics