Free 2-week DevOps health check for new engagements — pipelines, cloud spend and security reviewed. Claim yours

Automation with numbers attached

Ten engagements, what was actually wrong, what we built, and what changed afterwards. Client names are used with permission; where a name is withheld the sector and numbers are unchanged.

320+
Pipelines delivered
41%
Average cloud saving
68%
Faster release cycles
99.99%
Uptime achieved

AI data analyst for real-time conversational insights

Business Intelligence · Chatwit · Generative AI

The problem

Chatwit needed a more intuitive way for users to interact with business data and obtain insights without depending on SQL expertise or predefined dashboards. Traditional analytics workflows made new or exploratory questions dependent on technical teams.

What we built

  • Conversational AI data analyst using Amazon Bedrock and Amazon Nova Pro
  • Autonomous database exploration, schema discovery and database-specific SQL generation
  • Real-time AI processing using Amazon Bedrock ConverseStream and Socket.IO
  • Server-side SQL Guard allowing only controlled read operations and rejecting write or schema-changing statements
  • React and TypeScript frontend with Node.js and Express backend on Amazon ECS with AWS Fargate
  • Secure architecture using AWS IAM, AWS Secrets Manager, private VPC connectivity and DynamoDB conversation memory

The outcome

Chatwit gained an always-available AI data analyst that lets users ask questions naturally, watch the AI investigate the data, and receive results as metrics, tables or visualizations without needing SQL knowledge. The reusable query-agent also supports both MySQL and PostgreSQL environments.

Read full case study

GenAI-powered data exploration & self-service analytics

Self-Service Analytics · TNBT · Generative AI

The problem

TNBT had valuable operational data stored in relational databases, but accessing that information often required technical knowledge of database schemas and SQL. Business users needed answers to new and exploratory questions without depending on developers, analysts or database specialists.

What we built

  • Conversational data exploration platform using Amazon Bedrock and Amazon Nova Pro
  • Autonomous schema discovery, SQL generation and multi-step database investigation
  • Server-side SQL Guard allowing only controlled read operations and rejecting write or schema-changing statements
  • React and TypeScript frontend with Node.js and Express backend on Amazon ECS with AWS Fargate
  • Secure architecture using AWS Secrets Manager, IAM roles, private VPC subnets and DynamoDB conversation memory
  • Real-time agent processing through Socket.IO and WebSockets, with support for MySQL and PostgreSQL

The outcome

TNBT moved from a traditional reporting model toward self-service, conversational data exploration. Users can ask business questions in natural language while the GenAI agent explores database structures, generates and validates SQL, executes controlled read-only queries, handles errors, and presents insights as tables, metrics or visualizations.

Read full case study

AI-powered real estate intelligence through natural language

Real Estate · Trueestate · Generative AI

The problem

Business teams needed a faster and more accessible way to understand data stored across operational databases. Everyday questions about users, properties, activities, products, transactions and business performance often required knowledge of database structures and SQL, creating dependency on technical teams.

What we built

  • Conversational AI data analyst using Amazon Bedrock and Amazon Nova Pro
  • Agentic tool use to discover schemas, inspect columns and values, generate SQL and analyse results
  • Controlled server-side SQL validation with read-only operations and automatic rejection of write or schema-changing statements
  • AWS ECS Fargate application with Amazon RDS for MySQL and Amazon Aurora PostgreSQL
  • Secure operational controls using IAM, AWS Secrets Manager, private VPC subnets and Amazon DynamoDB conversation history

The outcome

Trueestate's data-access experience was transformed into a conversational GenAI workflow. Users can ask business questions in natural language and receive data-driven answers without knowing SQL or the underlying database schema.

Read full case study

Conversational business intelligence with autonomous SQL analytics

Business Intelligence · Dexlyn · Generative AI

The problem

Dexlyn's business teams needed timely answers from operational data, but accessing those insights often depended on people with SQL and database expertise. Traditional reporting workflows created a dependency on technical teams and were less effective when users had new questions that had not been anticipated in dashboards.

What we built

  • Conversational AI analytics agent using Amazon Bedrock and Amazon Nova Pro
  • Autonomous schema exploration, database-aware SQL generation and multi-step tool execution
  • Server-side SQL validation allowing only controlled read operations and rejecting write or schema-changing statements
  • React and TypeScript frontend delivered through Amazon S3 and Amazon CloudFront, with agent orchestration on ECS Fargate
  • Secure architecture using AWS IAM, AWS Secrets Manager, private VPC subnets and Amazon DynamoDB conversation context
  • Real-time agent processing through Socket.IO and WebSockets, with support for MySQL and PostgreSQL

The outcome

Dexlyn's traditional business intelligence workflow was transformed into a conversational, autonomous SQL analytics platform. Users can ask business questions in natural language while the AI agent explores the database, generates and validates SQL, handles errors, and presents business-friendly results.

Read full case study

Payments platform to zero-downtime releases

FinTech · Northwind Pay

The problem

A card-issuing platform released once a fortnight, inside a Saturday-night maintenance window, with a manual change-approval board and a rollback plan that consisted of a database backup and hope. Two of the previous six releases had overrun the window.

What we built

  • Rebuilt CI on GitHub Actions with reusable workflows shared across 40 services
  • Argo CD GitOps promotion across dev, staging and production with signed artefacts
  • Canary releases with automated rollback driven by error-rate and latency SLOs
  • Approval and segregation-of-duties controls enforced in the pipeline, generating audit evidence automatically

The outcome

Releases moved from fortnightly to on-merge within four months. The change-approval board still exists, but it reviews a generated report instead of blocking the release.

Black Friday scale at 40% lower cloud cost

Retail · Corevo Retail

The problem

The platform was provisioned year-round for its December peak. Cloud spend had grown 3× in two years while traffic grew 40%, and no team could say which line items were theirs — roughly a third of the bill was untagged.

What we built

  • Full cost attribution: a tagging standard enforced by policy, backfilled across the estate
  • Karpenter-based node autoscaling with spot pools for stateless workloads
  • Rightsized databases and caches, storage lifecycle policies and log retention tuning
  • Savings-plan and reserved-capacity model matched to the new steady-state baseline
  • Pre-peak load testing and a game day rehearsing the failure modes

The outcome

Steady-state spend fell 40% while peak capacity headroom doubled. The following Black Friday ran at 99.99% availability with no manual scaling intervention.

HIPAA-ready platform in eleven weeks

HealthTech · Vantiq Health

The problem

A Series A health startup had a single AWS account, shared credentials, no infrastructure code and an enterprise customer requiring evidence of HIPAA controls within one quarter. Their engineers were spending a day a week on manual compliance screenshots.

What we built

  • Multi-account landing zone with SSO, guardrails and separated production data
  • Every resource in Terraform modules, environments reproducible from scratch
  • Encryption in transit and at rest, key rotation and PHI-safe non-production data masking
  • Kyverno admission policy plus continuous control monitoring mapped to HIPAA safeguards
  • Automated evidence collection feeding their compliance platform

The outcome

Audit-ready in eleven weeks, the enterprise contract signed on schedule, and manual control work reduced by 85%.

Kubernetes migration without a visible outage

Gaming · Skyforge Games

The problem

Matchmaking and live-ops services ran on hand-configured virtual machines across three regions. Scaling for a launch took two days of lead time, on-call was paged most nights, and no two environments were configured the same way.

What we built

  • Three regional EKS clusters defined by a single Terraform module
  • Progressive workload migration behind weighted DNS, one service at a time
  • Autoscaling tied to matchmaking queue depth, with session draining on scale-down
  • Prometheus and Grafana with SLOs per game mode, and alerts rewritten around symptoms

The outcome

The migration completed with no customer-visible downtime. Out-of-hours pages dropped by two thirds, and a regional capacity increase now takes minutes.

One platform for twelve product teams

SaaS · Lumen Labs

The problem

Twelve teams had built twelve slightly different pipelines. Standing up a new service took three weeks of copy-paste, monitoring was inconsistent, and nobody could answer who owned a given service at 3am.

What we built

  • A Backstage developer portal with an authoritative service catalogue and ownership records
  • Golden path templates: a new service ships with CI, dashboards, alerts and on-call routing on day one
  • Self-service ephemeral environments provisioned per pull request
  • Maturity scorecards showing each team where they diverge from the paved road

The outcome

New service setup fell from three weeks to under a day, and eleven of twelve teams adopted the golden path within two quarters — without it ever being mandated.

From nightly firefighting to a quiet pager

Logistics · Orbit Logistics

The problem

A dispatch platform with 4,000 alert rules, most of them noise. Mean time to restore was over four hours because the first hour was always spent working out which alert mattered. Two engineers had resigned citing on-call.

What we built

  • OpenTelemetry tracing across the dispatch path, replacing guesswork with evidence
  • SLOs per customer journey, with alerting rebuilt around symptoms rather than causes
  • Alert rules cut from 4,000 to 180, each with an owner and a runbook link
  • Incident severity model, comms templates and blameless postmortems with tracked actions
  • Managed 24×7 cover during the transition while the team rebuilt confidence

The outcome

Mean time to restore fell from over four hours to 38 minutes, and weekly pages dropped from 31 to 4.

Client feedback

In their words

We went from a two-week release train to deploying on merge. Grey Bracket rebuilt our pipelines in our own repositories and taught the team as they went — nothing was a black box at the end of it.

RMR. MehtaVP Engineering, Northwind Pay

Our AWS bill had grown faster than our revenue. The FinOps review paid for the whole engagement inside the first quarter, and the tagging model means it has stayed down since.

SOS. OkonkwoCTO, Corevo Retail

They took our SOC 2 evidence collection from a month of screenshots to something that generates itself out of the pipeline. Our auditors had fewer questions than in any previous year.

JLJ. LindqvistHead of Platform, Vantiq Health

The Kubernetes migration happened without a single customer-visible outage, and the on-call rota is quieter than it was on the old VMs. That is the part I did not expect.

AKA. KaurDirector of Engineering, Skyforge Games

Your project could be the next one on this page

Start with the free two-week health check. You will get the same baseline analysis that every engagement on this page began with.