Flow observability for data pipelines and services

Know which part broke, and why, in minutes.

Draw your system: databases, proxies, Kafka, queues, APIs. MDJaada checks every part, and when something fails it tells you the likely cause, what changed upstream, and shows the evidence.

  • Self-hosted: your data stays in your network
  • Running in 30 to 45 minutes
  • No code changes needed to start
Checkout pipeline
MDJaada map of a checkout pipeline: orders-db is down and marked as the number one likely cause; pgbouncer and orders-api are down as symptoms; Kafka topic and billing consumer group are degraded.

“orders-api is down.”

Is it the API? The connection pool? The database? The migration someone ran at 05:16? Kafka? Most alerts tell you where it hurts. Your team then spends the first half hour of every incident working out where it started. MDJaada does that part for you.

How it works

Three steps from install to answers

  1. 1

    Install on one server

    Unpack the package on a Linux server with Docker and run ./setup.sh. A small agent checks your systems from inside your network. It only connects outwards. Nothing needs to reach in.

  2. 2

    Draw your flow

    Drag your databases, proxies, queues and services onto the canvas and connect them in the direction data flows. Kafka topics, consumer groups, RabbitMQ queues and Kubernetes services are discovered for you.

  3. 3

    Get the cause, not just the alert

    When a part fails, MDJaada opens one incident, ranks the likely causes with the reasons behind each, marks the symptoms, and alerts you in Slack, PagerDuty or email.

Incidents

One incident, ranked causes, the evidence

  • Likely causes, ranked. Each with its reasons: first to fail, most upstream, changed just before. Downstream failures are marked as symptoms, so you get one alert instead of ten.
  • What changed. Deploys, migrations and Kubernetes rollouts recorded next to the failure: “schema change at 05:16: migration 0042”.
  • Evidence from your telemetry. Error spans and log lines from the failure window, grouped, with the full trace one click away. Queried inside your network at the moment it is needed.
  • Metrics around the failure. Charts of the suspects from 15 minutes before, with the moment of failure marked.
Incident panel: orders-db failed; likely causes ranked orders-db first with a schema change at 05:16, then pgbouncer and orders-api as probable symptoms; data-age chart with the failure marked.
Kafka consumer group billing: degraded, lag 3339 above the warn line of 1000, with charts of lag, messages consumed per second and members.

Queues, streams and metrics

Producer side or consumer side? It says which.

  • Kafka as layers. Cluster, topics and consumer groups as separate nodes. A lost broker is reported on the cluster, “no new messages” on the topic (the producer), lag and stalls on the group (the consumer). RabbitMQ and Pub/Sub work the same way.
  • Charts on every node. Lag, data age, latency, queue depth: every number is charted for 7 days with your warn and fail lines, and a small trend line sits on each node on the map.
  • Your apps' own metrics. Send OpenTelemetry metrics and judge any node on them: p95 latency, error rate in %, pool usage, orders per minute.
  • Works with Grafana. A Prometheus endpoint exposes node states, numbers and incidents for your existing dashboards.

Integrations

Checks for the parts your data flows through

Every node type has checks that know what “broken” means for it, from replication lag to consumer lag to a DAG that stopped running.

Databases and stores

  • Postgres
  • MySQL / MariaDB
  • ClickHouse
  • Snowflake
  • MongoDB
  • Redis
  • Elasticsearch / OpenSearch

Streams and queues

  • Kafka clusters
  • Kafka topics
  • Kafka consumer groups
  • RabbitMQ
  • AWS SQS
  • Google Pub/Sub

Services and platforms

  • HTTP services and APIs
  • Proxies (nginx, PgBouncer, Envoy)
  • Kubernetes workloads
  • AWS CloudWatch
  • Azure Monitor

Pipelines, telemetry, alerts

  • Airflow (2 and 3)
  • MLflow models
  • OpenTelemetry traces, logs, metrics
  • Prometheus metrics
  • Slack, PagerDuty, email

Install

Up and running in under an hour

You receive the install package from us when the pilot starts. Everything runs on your own servers; nothing is installed on your laptops or inside your applications.

1

The MDJaada server

One Linux server or VM with Docker · 2 CPU, 4 GB RAM

tar xzf mdjaada-selfhost-<version>.tar.gz
cd mdjaada-selfhost-<version>
./setup.sh

Prints the web address and the login. An agent runs on this server too, so systems it can reach are checked straight away.

2

Agents in other networks, no Docker or Python needed

Linux with systemd (x86-64 or ARM64) · outbound access to the MDJaada server

curl -fsSLO https://github.com/madhubolla2205/mdjaada-releases/releases/latest/download/install.sh
sudo bash install.sh --api https://<your-mdjaada-server>:8080/api \
  --agent-id dc1-agent --assigned-only
  • Downloads a single-file agent (about 28 MB, database and Kafka drivers inside) and verifies its checksum and signature before installing.
  • Asks for the API key, so it never lands in your shell history.
  • Starts on boot and restarts on failure, as its own user with a read-only view of the system.
  • Passwords for your databases stay in /etc/mdjaada-agent/secrets.env on that server.
  • In MDJaada, set Checked by agent to dc1-agent on the nodes it should check.
journalctl -u mdjaada-agent -f                    # agent logs
sudo bash install.sh --upgrade                    # newer version (with the newer install.sh), settings kept
sudo bash install.sh --uninstall                  # remove (--purge also deletes its data)

Runs on RHEL/Alma/Rocky 8+, Ubuntu 20.04+ and Debian 10+. Checks AWS, Google Cloud, Azure or Snowflake? Use the Python install from the package instead. Prefer containers? The agent also runs as a Docker container, or inside Kubernetes with read-only permissions. Apps can send OpenTelemetry to any agent on port 4318. Releases: github.com/madhubolla2205/mdjaada-releases.

Security and data

Built to run inside your network

Self-hosted

MDJaada runs on your own server. Status, incidents, traces and logs stay there. Nothing is sent to us.

Outbound-only agent

Agents in other networks or Kubernetes clusters connect out to your MDJaada server. No inbound ports to open.

Secrets stay on the host

Passwords live in a file on the agent's server and are referenced as env:NAME. They never enter MDJaada's database.

Read-only

Checks only read: a monitoring user on the database, describe rights on Kafka, read-only RBAC in Kubernetes.

Free pilot

Try it on your own system for 2 to 4 weeks

  1. A 30-minute call. You tell us which systems hurt most; we agree what success looks like.
  2. Install together. One Linux server with Docker, read-only users for your databases and Kafka. About 30 to 45 minutes to a live map.
  3. Run it through real incidents. If it finds the cause faster, we talk about continuing. If not, you remove it with one command.

Request a free pilot

What do you run? (optional)

We use these details only to reply about your pilot. Privacy

FAQ

Questions teams ask

Do we need to change our code?

No. Health checks, database queries, Kafka lag and the rest work without touching your applications. If your apps already send OpenTelemetry, point them at MDJaada to add traces, log lines and app metrics to incidents.

What does it need to run?

One Linux server or VM with Docker (2 CPU, 4 GB RAM is enough for a pilot) that can reach the systems you want to watch. Systems in other networks or clusters get their own small agent: on a plain Linux server it installs as a systemd service from a single signed file, with no Docker or Python needed. See Install.

Does it replace Grafana, Datadog or our APM?

No. Those show you everything; MDJaada follows the path between your systems and points at where a failure started. It sits next to them, and exposes a Prometheus endpoint so its findings can appear in Grafana.

Where is our data stored?

On your server. Traces, logs and app metrics are kept by the agent inside your network and only queried there when an incident needs them.

How do alerts reach us?

Slack, PagerDuty or email. Incidents can be acknowledged, nodes muted for an hour or a day, and maintenance windows scheduled so planned work does not page anyone.

What does it cost?

The pilot is free. If you want to keep using it afterwards, we agree a price based on the size of your system.

Spend incidents fixing, not searching.

Request a free pilot