PCA sample questions with answers
10 free PCA sample questions across the exam's domains, each with its answer and an explanation. No account needed.
PCA sample questions
Prometheus Certified Associate (PCA), CNCF / The Linux Foundation.
Question 1
Domain: Observability Concepts
What distinguishes a service level agreement from a service level objective?
- An agreement is internal to the engineering team while an objective is published to users
- An agreement is measured automatically while an objective is assessed by hand
- An agreement attaches explicit consequences, such as a rebate, to missing its objectives
- An agreement always covers availability while an objective always covers latency
Show the answer
Answer: C. An agreement attaches explicit consequences, such as a rebate, to missing its objectives
The simple test is to ask what happens if the target is missed: with no explicit consequence it is an objective, not an agreement. The internal-versus-published answer inverts normal practice, since teams often keep a tighter internal objective than the figure advertised to users.
Checked against: https://sre.google/sre-book/service-level-objectives/
Question 2
Domain: Observability Concepts
What does a service discovery mechanism provide to a Prometheus scrape configuration?
- An automatically refreshed target list, with __meta_ labels describing each object
- A fixed set of targets that is read once when Prometheus starts
- A copy of the metrics exposed by each target, so Prometheus does not have to scrape them
- An authentication token for every target in the environment
Show the answer
Answer: A. An automatically refreshed target list, with __meta_ labels describing each object
Service discovery keeps the target list current and attaches __meta_ labels that relabelling rules can use to filter targets and shape their label sets. A fixed set of targets read once at start-up describes static_configs, which is the alternative to service discovery rather than a description of it.
Checked against: https://prometheus.io/docs/prometheus/latest/configuration/configuration/#scrape_config
Question 3
Domain: Prometheus Fundamentals
A finance team proposes using the existing Prometheus deployment to count billable API calls per customer, and to invoice directly from the resulting series.
How should the monitoring team respond?
- Accept, but set the scrape interval to one second so that no request is missed
- Decline, because Prometheus cannot store counters that grow beyond 32 bits
- Decline: Prometheus is not accurate enough for per-request billing; use a separate billing system
- Accept, but add a customer label to every request counter so the data can be split later by account
Show the answer
Answer: C. Decline: Prometheus is not accurate enough for per-request billing; use a separate billing system
The documentation is explicit that Prometheus is unsuitable where 100 per cent accuracy is required, such as per-request billing, because sampling and missed scrapes make the data incomplete. Shortening the scrape interval is the tempting engineering fix, but sampling a counter more often does not make the collected data exact, and it multiplies storage and load.
Checked against: https://prometheus.io/docs/introduction/overview/#when-does-it-not-fit
Question 4
Domain: Prometheus Fundamentals
A team must keep alerting available while one monitoring host is patched. They are worried that duplicating the setup will send every page twice.
Which arrangement meets both requirements?
- Run two Prometheus servers with the rule files split between them so no alert is defined twice
- Run one Prometheus server and take a snapshot before patching
- Run one Prometheus server and two Alertmanagers behind a load balancer
- Two identical Prometheus servers sending to a clustered Alertmanager that deduplicates alerts
Show the answer
Answer: D. Two identical Prometheus servers sending to a clustered Alertmanager that deduplicates alerts
Redundancy in Prometheus means running identical servers; identical alerts are deduplicated by the Alertmanager, and Alertmanager instances can be clustered for their own availability. Splitting the rule files is the closest wrong answer: it does avoid duplicate pages, but each alert then has a single point of failure, which defeats the purpose.
Checked against: https://prometheus.io/docs/introduction/faq/#can-prometheus-be-made-highly-available
Question 5
Domain: PromQL
The join rate(http_requests_total[5m]) * on (instance) group_left (version) app_build_info worked for months but now fails during rollouts with an error about duplicate series for the match group.
What is the most likely cause?
- The rate() range is shorter than the scrape interval during rollouts
- on (instance) must list every label that the two sides share, including job
- group_left cannot be used when the left-hand side has more than one series for any single instance
- During rollouts an instance briefly has two app_build_info series, so the "one" side is not unique
Show the answer
Answer: D. During rollouts an instance briefly has two app_build_info series, so the "one" side is not unique
In a many-to-one match every match group must have exactly one element on the "one" side; two build_info series for the same instance (for example old and new version within the lookback window) break that rule and the query errors. Having many series on the left is precisely what group_left permits, so that option inverts the requirement.
Checked against: https://prometheus.io/docs/prometheus/latest/querying/operators/#vector-matching
Question 6
Domain: PromQL
A nightly batch job pushes my_job_last_success_timestamp_seconds to the Pushgateway after each successful run.
Which alert expression fires when the job has not succeeded for more than 25 hours?
- changes(my_job_last_success_timestamp_seconds[25h]) > 0
- time() - my_job_last_success_timestamp_seconds > 25 * 3600
- my_job_last_success_timestamp_seconds > 25 * 3600
- absent(my_job_last_success_timestamp_seconds) > 25 * 3600
Show the answer
Answer: B. time() - my_job_last_success_timestamp_seconds > 25 * 3600
Exporting the timestamp of the last success and comparing it to time() is the recommended pattern for batch jobs; the difference is the age of the last success. The changes() option inverts the logic: it is true when the job did succeed in the window, so it would fire on the healthy case.
Checked against: https://prometheus.io/docs/practices/instrumentation/#batch-jobs
Question 7
Domain: PromQL
Each of four instances exposes request_duration_seconds_sum and _count. One instance serves 90% of traffic. A dashboard shows avg(rate(request_duration_seconds_sum[5m]) / rate(request_duration_seconds_count[5m])) as the service's average latency.
What is wrong and what should replace it?
- It weights instances equally; divide sum(rate(..._sum[5m])) by sum(rate(..._count[5m]))
- rate() cannot be applied to _sum; use increase() for the numerator
- It must use max() rather than avg() because latencies cannot be averaged
- Nothing; averaging the per-instance averages always equals the overall request-weighted average
Show the answer
Answer: A. It weights instances equally; divide sum(rate(..._sum[5m])) by sum(rate(..._count[5m]))
Averaging per-instance means gives a low-traffic instance the same weight as the busiest one; summing numerators and denominators before dividing yields the true request-weighted mean. The _sum series is a counter, so rate() is correct on it, and there is no rule against averaging latencies when it is done this way.
Checked against: https://prometheus.io/docs/practices/histograms/
Question 8
Domain: Instrumentation and Exporters
A service counts failed payment calls with payment_failures_total. An analyst asks what fraction of payment calls fail.
What should the service also expose?
- A log line for every success, so the ratio can be derived by parsing logs
- A counter of total payment attempts, so failures can be divided by attempts
- Nothing more; the ratio can be derived from payment_failures_total alone
- A gauge set to the current failure percentage, calculated in the application
Show the answer
Answer: B. A counter of total payment attempts, so failures can be divided by attempts
When reporting failures you should also have a metric for the total number of attempts, which makes the failure ratio easy to compute with rate() on both. A gauge holding a percentage computed in the application loses the ability to aggregate correctly across instances and over arbitrary windows.
Checked against: https://prometheus.io/docs/practices/instrumentation/#failures
Question 9
Domain: Alerting & Dashboarding
Three Alertmanager replicas are run for high availability. An engineer puts them behind a load balancer and points every Prometheus server at the load balancer address.
What is the recommended setup instead?
- Point each Prometheus at a different single Alertmanager to split the load
- List every Alertmanager in each Prometheus and cluster the Alertmanagers for deduplication
- Run only one Alertmanager, because Alertmanager cannot be made highly available in any supported way
- Keep the load balancer but enable sticky sessions per alert
Show the answer
Answer: B. List every Alertmanager in each Prometheus and cluster the Alertmanagers for deduplication
Prometheus should send alerts to all Alertmanagers directly, and the Alertmanager cluster uses gossip to deduplicate so each notification goes out once; load balancing between them is explicitly discouraged. Assigning each Prometheus a single Alertmanager defeats the purpose, since that Alertmanager becomes a single point of failure for its alerts.
Checked against: https://prometheus.io/docs/alerting/latest/alertmanager/#high-availability
Question 10
Domain: Alerting & Dashboarding
When a cluster's ClusterUnreachable alert (severity critical) is firing, every warning alert from that cluster is noise.
Which Alertmanager configuration suppresses those warnings?
- A silence defined at start-up in the Alertmanager configuration file for severity="warning"
- group_by: [cluster] on the warning route
- An inhibit rule: source severity="critical", target severity="warning", equal: [cluster]
- A route with continue: true for severity="warning"
Show the answer
Answer: C. An inhibit rule: source severity="critical", target severity="warning", equal: [cluster]
Inhibition mutes target alerts while a matching source alert fires, and equal restricts it to alerts sharing the listed labels, here the same cluster. Silences are created at runtime through the UI or API rather than in the configuration file, and they are time-bound rather than conditional on another alert.
Checked against: https://prometheus.io/docs/alerting/latest/configuration/#inhibit_rule
More practice
A 20-question practice sampler is free with an account; Pro adds the full question bank and timed mock exams.
PCA course and practice exam: Prometheus Certified Associate: the exam guide, with the format, cost, pass mark and domains from the vendor.