CommandLink

Staff Software Engineer, AI Investigation & Triage

CommandLink  •  Colombia, CO (Remote)  •  14 hours ago
Apply
AI can make mistakes so check important info. Chat history is never stored.

Job Description

About Command|Link

Command|Link is a global SaaS Platform providing network, voice services, and IT security solutions, helping corporations consolidate their core infrastructure into a single vendor and layering on a proprietary single pane of glass platform. Command|Link has revolutionized the IT industry by tackling the problems our competitors create. In recognition for our unprecedented innovation and dedication, Command|Link was recognized as the SD-WAN Product of the Year, ITSM Visionary Spotlight, UCaaS Product of the Year, NaaS Product of the Year, Supplier of the Year, and the AT&T Strategic Growth Partner. Command|Link has built the only IT platform for scale that solves ISP vendor sprawl and IT headaches. We make it easy for our customers to get more done, maximize uptime and improve the bottom line.

Learn more about us here!

This is a 100% remote position

About your new role:

This role sets the technical direction for the AI-powered layer that sits on top of our alert engine and dramatically reduces mean time to triage. You'll define how LLM reasoning and graph context combine to group alerts into incidents, infer likely narratives, and brief analysts, and you'll make sure that architecture holds up as it's extended across the teams that feed it: alerting and detection, graph and topology, and the analyst-facing investigation experience.

Our edge is correlation: taking diverse sources such as security tooling, monitoring telemetry, syslog, OpenTelemetry, and L2-L4 network protocols, deriving usable network and system topologies and the dependencies between them, and giving that context to LLMs to reason over for correctness, troubleshooting, remediation, and investigation. You'll be the person other engineers come to when a correlation, reasoning, or workflow-durability problem doesn't have an obvious answer yet.

Key Responsibilities:

  • Facilitate the architecture decisions that shape how alert correlation, LLM reasoning, and graph context combine across the investigation and triage pipeline, decisions that other teams (alerting, graph/topology, analyst UX) build on.
  • Own the design for how incidents get grouped, how confidence scores and priority tiers get assigned, and how the AI infers and presents likely incident narratives to SOC/NOC analysts, holding that design to a bar that measurably drives down mean time to triage.
  • Set the technical standard for the collaborative investigation experience: comments, side-bar conversations, shared troubleshooting, and AI-initiated ticketing as a post-triage action.
  • Define the durable workflow architecture, proto-first contracts, deterministic execution, and versioning discipline that this system and adjacent teams' workflows are built on.
  • Step into other teams or projects when a triage-latency, narrative-accuracy, or workflow-durability problem is stuck, contributing hands-on code and design across the alerting, graph-context, and analyst-workbook components as needed, not just within your own team.
  • Balance near-term delivery of MTTT-reducing capability against the long-term technical foundation (workflow durability, model routing, graph schema) that the rest of the org will build on.
  • Actively mentor engineers working across these teams and represent this architecture in conversations with stakeholders outside engineering.
  • Takes on additional responsibilities and projects as needed to support the success of the team and organization.

What you'll need for success:

Required

  • Recognized as an authority in Go and Python, with a track record of building production systems others rely on.
  • Deep, hands-on expertise in LLM tool calling and model routing, and in applying LLM reasoning to structured, real-world problems rather than prototypes.
  • Mastery of Temporal workflow design: determinism, versioning, and proto-first contracts, with experience making these decisions for systems other teams depend on.
  • Demonstrated ability to build novel solutions where no existing playbook applies, ideally including systems that combine graph-based reasoning (e.g., Memgraph or similar) with LLM-driven decision-making.
  • Experience with multi-channel, real-time systems, and comfort operating across a broader technical footprint: Kubernetes/Helm/Docker, AWS/Azure/GCP, Kafka, OpenSearch, and infrastructure-as-code (Argo, Spacelift).
  • A deep, working understanding of the telemetry and protocols this system reasons over, including OpenTelemetry, syslog, SNMP, NetFlow/sFlow, ICMP, and firewall logs, and the judgment to turn that data into usable network and system topologies.
  • Track record of influencing architecture decisions across multiple teams, not just within a single codebase.

Nice to Have

  • Experience with stream processing (e.g., Flink) or detection/anomaly-detection systems (e.g., OpenSearch alerting).
  • Familiarity with endpoint and cloud inventory tooling (osquery, Steampipe) or telemetry pipelines (Vector).
  • Prior experience representing engineering work externally, at conferences, in customer conversations, or in published writing.

Why you'll love life at Command|Link

Join us at CommandLink, where you'll have the opportunity to shape the future of business communication. We value the innovative spirit and seek individuals ready to bring their unique vision and expertise to a team that values bold ideas and strategic thinking. Are you ready to make an impact?

  • Room to grow at a high-growth company
  • An environment that celebrates ideas and innovation
  • Your work will have a tangible impact
  • Flexible time off
  • Fun events at cool locations
  • Employee referral bonuses to encourage the addition of great new people to the team

At CommandLink, we’re committed to creating a fair, consistent, and efficient hiring experience. As part of our process, we use AI-assisted tools to help review and analyze applications. These tools support our recruiting team by identifying qualifications and experience that align with the requirements of each role.

AI tools are used only to assist in the evaluation process — they do not make final hiring decisions. Every application is reviewed by a member of our recruiting or hiring team before any decisions are made.

CommandLink

About CommandLink

CommandLink is the only global infrastructure provider that unifies connectivity, security, voice, and AI into one intelligent platform—backed by world-class support.

At the foundation of our platform is the industry’s most advanced Global ISP Aggregation fabric, giving enterprises access to 5,000+ carriers in over 200 countries through a single vendor, contract, and software interface. Whether you're deploying a single site or a global hybrid network, CommandLink delivers unmatched reach, control, and SLA-backed performance.

We integrate SD-WAN, SASE (Fortinet, Versa, Cato, Meraki, Juniper), MDR/XDR, Cloud Voice, monitoring, and support into one seamless experience—purpose-built for modern IT teams managing multi-vendor environments.

CommandLink’s platform includes AI-driven engines, Command|Monitor and Command|Alert, that proactively detect network degradation, security threats, or service anomalies and automatically trigger intelligent workflows, vendor escalations, or internal remediation—before users are affected.

Our Command|POD support model sets a new standard in enterprise service. Each customer is assigned a dedicated team of Tier 3 engineers who understand your environment and provide real-time, proactive support. No call centers. No handoffs. No excuses.

From first mile to last, CommandLink consolidates and simplifies global IT infrastructure into a single platform—with intelligent operations and unmatched support.

Learn more at www.CommandLink.com or contact sales@commandlink.com

Industry
IT & Software
Company Size
201-500 employees
Headquarters
Bothell, WA
Year Founded
2012
Social Media