Sign In Get Started
Full Time Remote

Systems Observability Specialist

📍 Remote / Worldwide, United States 💵 $ 135,000 – 155,000 / yr 21 hours ago

Department: Accounting & Finance Shift: Flexible Shift Location: Remote Salary: $ 135,000 – 155,000 / yr Application Deadline: October 29, 2026

Bright Vision Technologies is a technology consulting and software development company delivering advanced cloud, artificial intelligence, data, and enterprise technology solutions to organizations across the United States.

The company provides an opportunity for experienced technology professionals to contribute to complex projects while developing their careers within an established and growing technology organization.

About the Role

Bright Vision Technologies is seeking an experienced Systems Observability Specialist to design, implement, and operate the observability infrastructure that helps engineering teams understand the health, reliability, and performance of their systems.

This position will work across the complete observability ecosystem, including metrics, logging, distributed tracing, monitoring, dashboards, telemetry pipelines, storage, and alerting.

You will be responsible for building observability solutions that transform large volumes of system telemetry into meaningful and actionable information.

The goal is to provide engineering teams with reliable signals that help them detect problems faster, understand system behavior, improve performance, and operate production environments with greater confidence.

The successful candidate should have experience developing and operating observability platforms at scale and understand the advantages and trade-offs associated with both open-source and SaaS-based monitoring solutions.

Visa and Work Authorization

Bright Vision Technologies welcomes applications from U.S. Citizens, Green Card holders, EAD holders, and candidates eligible for an H-1B transfer.

Please note that the company is unable to sponsor new H-1B visa petitions for this particular position.

In this role, you will help establish and maintain reliable observability capabilities across complex production environments. Responsibilities may include:

  • Design and operate platforms for metrics collection, logging, distributed tracing, monitoring, and alerting.
  • Develop and maintain telemetry collection agents, pipelines, dashboards, storage systems, and alerting workflows.
  • Improve the quality of monitoring signals while reducing unnecessary noise and ineffective alerts.
  • Build observability solutions that provide useful information to engineering teams and relevant business stakeholders.
  • Help teams identify system failures, performance issues, and operational risks more efficiently.
  • Evaluate open-source and commercial observability technologies based on scalability, reliability, usability, cost, and organizational requirements.
  • Support high-throughput and high-cardinality metrics and logging environments.
  • Integrate observability capabilities with CI/CD pipelines, incident-management platforms, and engineering workflows.
  • Help establish appropriate Service Level Objectives (SLOs), error budgets, and reliability practices.
  • Identify opportunities to improve the operational and financial efficiency of observability infrastructure.

Requirements

Candidates should have a Bachelor's degree in Computer Science or a related technical discipline, along with substantial professional experience in reliability, infrastructure, or observability engineering.

The ideal candidate should also have:

  • 5+ years of professional experience in Site Reliability Engineering (SRE), platform engineering, observability, or a closely related discipline.
  • Advanced hands-on experience with Prometheus and Grafana.
  • Experience with at least one major commercial observability platform such as Datadog, New Relic, or Splunk.
  • Strong understanding of OpenTelemetry, distributed tracing, and structured logging.
  • Programming experience with at least one general-purpose language such as Go, Python, or Java.
  • Experience managing high-cardinality and high-throughput metrics and logging pipelines in production environments.
  • Strong understanding of SRE principles, SLOs, SLIs, and error budgets.
  • Experience integrating monitoring and observability systems with CI/CD and incident-management tools.
  • Solid understanding of Linux internals, networking concepts, containers, and containerized platforms.
  • Excellent communication and collaboration skills, including the ability to explain technical information clearly to different audiences.

Preferred Qualifications

Additional experience that may strengthen your application includes:

  • Hands-on experience operating technologies such as Thanos, Grafana Mimir, Cortex, Loki, or Tempo at scale.
  • Contributions to OpenTelemetry or other open-source observability projects.
  • Familiarity with eBPF-based observability and monitoring technologies.
  • Experience developing and implementing strategies to reduce observability infrastructure and telemetry costs.
  • Experience working within regulated industries or environments requiring audit-grade logging, data retention, and compliance controls.

Ideal Candidate

This opportunity is particularly suited to an experienced observability or platform engineer who understands that effective monitoring involves more than collecting large amounts of telemetry.

The successful candidate will be able to determine which signals matter, design systems that make those signals accessible, and help engineering teams convert telemetry into practical insights that improve system reliability and operational decision-making.

You should be comfortable working with large-scale production systems, evaluating different technical approaches, collaborating across engineering teams, and balancing performance, reliability, usability, and cost.