Big Data Prep https://www.bigdataprep.com/ Make Big Data Easier To Use Tue, 29 Sep 2026 12:43:13 +0000 en-US hourly 1 https://wordpress.org/?v=7.1.2 https://www.bigdataprep.com/wp-content/uploads/2021/12/cropped-BigDataPrep-Mini-Logo-32x32.png Big Data Prep https://www.bigdataprep.com/ 32 32 NCP-AI Prerequisites: Is the Nutanix AI Exam for You? https://www.bigdataprep.com/2026/09/29/ncp-ai-prerequisites-is-the-nutanix-ai-exam-for-you/ Tue, 29 Sep 2026 00:00:00 +0000 https://www.bigdataprep.com/?p=16449 An infrastructure engineer weighing up an AI credential needs one answer first: is it written for me? Here is the experience Nutanix assumes, and a way to test yourself against it.

The post NCP-AI Prerequisites: Is the Nutanix AI Exam for You? appeared first on Big Data Prep.

]]>

You run clusters for a living, your organisation has started talking about private inference, and someone has suggested you sit NCP-AI. The decision in front of you is whether Nutanix Certified Professional – Artificial Intelligence, the Nutanix credential built around Nutanix Enterprise AI, is meant for an infrastructure engineer like you or for the data science team down the corridor.

It is meant for you, with conditions. Nutanix describes the successful candidate in unusually concrete terms: years of virtual infrastructure work, a year of cloud native and Linux command line experience, and Kubernetes knowledge at the level of a working cluster administrator. This article sets out the NCP-AI prerequisites one by one, shows where each appears in the exam objectives, and gives you a way to check your own readiness before you spend 200 USD on an attempt.

Table of Contents

  1. Who Is the NCP-AI Exam Actually Written For?
  2. What Are the NCP-AI Prerequisites Nutanix Names?
  3. Can You Answer These Before You Book?
  4. How Much Kubernetes Does CKA-Level Knowledge Mean?
  5. What You Do Not Need: Data Science and Model Training
  6. What Do You Commit To When You Book NCP-AI?
  7. NCP-AI, NCP-CN or NCP-MCI: Which Professional Exam Fits?
  8. How Do You Close the Gaps in the Right Order?
  9. Frequently Asked Questions
  10. Conclusion

Who Is the NCP-AI Exam Actually Written For?

The NCP-AI exam is written for infrastructure and platform engineers who will install, configure, optimise and troubleshoot Nutanix Enterprise AI, and connect generative AI applications and agents to it. It suits virtualisation administrators moving towards Kubernetes and GPU workloads. It is not aimed at data scientists, because no objective asks you to train or tune a model.

Who NCP-AI is meant for: cluster admins, platform teams and NAI consultants, against those with no Linux shell, no cluster time or model work only

That matters because the name misleads in both directions. Infrastructure engineers see “Artificial Intelligence” and assume the exam belongs to someone else. Machine learning practitioners see it and assume their modelling skills will carry them. Neither is right.

The job the exam describes

Read the five syllabus sections as a job description. You stand up the platform, you import a large language model, you publish it behind an endpoint, you hand an API key to an application team, and you work out what broke when the endpoint slows down. Every one of those tasks is operations work.

The Nutanix Enterprise AI overview makes the same point from the product side. The platform runs on Kubernetes, it needs GPUs, and it can sit on Nutanix Kubernetes Platform or on a managed service such as EKS, AKS or GKE. Whoever looks after that stack is the person the credential is describing.

Three readers, three answers

  • A Nutanix or VMware administrator with some Kubernetes exposure: yes, after closing the Kubernetes gap.
  • A platform engineer who already runs Kubernetes in production: yes, after learning the Nutanix layer and GPU sizing.
  • A data scientist with no infrastructure background: not yet, because the assumed experience is the hard part, not the AI vocabulary.

“To succeed in earning the NCP-AI certification, you should bring a strong foundation in virtualized infrastructure including virtual machines, hypervisors, and virtual networking.”

Suzanne DeWitt, Nutanix Community Education Blog

What Are the NCP-AI Prerequisites Nutanix Names?

The NCP-AI prerequisites are a candidate profile, not a list of required certificates. Nutanix says successful candidates have at least three years of virtual infrastructure experience and one year with cloud native technologies and the Linux command line, plus knowledge of NCI, cloud IaaS, GPUs, Nutanix Unified Storage and Kubernetes at Certified Kubernetes Administrator level.

The official NCP-AI page sets this out for version 6.10 of the exam, which awards the NCP-AI 6 certification. What the page does not do is tell you why each item is there. The objectives answer that.

What Nutanix expects Where it appears in the objectives What happens without it
3 years of virtual infrastructure DNS, FQDN and certificate setup; storage classes; infrastructure performance views Deployment questions read as unfamiliar networking puzzles
1 year of cloud native work NKP and non NKP installation; dark site installs; KServe as a prerequisite You cannot tell a platform fault from a product fault
Linux command line Querying endpoints with Python or Curl; reading tool calling commands The application section becomes guesswork
GPU knowledge Choosing the number and type of GPUs for a model; spotting an endpoint on CPU Sizing scenarios have no anchor
Nutanix Unified Storage and NCI Storage classes and CSI driver connectivity Model import failures look random
CKA-level Kubernetes System resources, taints, allocatable compute, container image failures The troubleshooting section is out of reach

Notice how little of that table is specific to AI. Five of the six rows would sit comfortably in any infrastructure exam. The AI content arrives on top of them, in model import, endpoint creation and output quality, and it assumes the base is already there.

Experience is not the same as eligibility

Nothing on the official page says you will be refused a booking without these years of experience. They describe who passes, not who may sit. Treat them as an honest forecast. If you are two years short on infrastructure and have never touched a cluster, the forecast is poor, however much you read about language models.

Can You Answer These Before You Book?

A quick readiness check for NCP-AI is to read real sample questions cold, before any study. If you can explain why each wrong option is wrong, your infrastructure base is sound. If the questions read as unfamiliar nouns, you have found your gap early and cheaply, before paying for an attempt.

The published samples are short scenarios. One describes a health check failing at the serving layer with KServe pods in error, and asks what breaks. Another describes an endpoint falling back to CPU acceleration and asks what to verify. A third gives you a certificate error where the FQDN is missing from the SAN field.

Work through the NCP-AI sample questions once with no preparation, then sort your misses into three piles.

  • Misses about DNS, ingress and certificates point to the infrastructure base.
  • Misses about pods, taints, labels and KServe point to Kubernetes.
  • Misses about endpoints, API keys, sample requests and inference engines point to the product itself.

The third pile is the easy one. Product knowledge comes from the recommended course and from time in the interface. The first two piles take longer, and they are the ones the stated experience is there to cover.

What a good result looks like

Do not read a score into ten questions. Read the pattern. An engineer who misses only product questions is weeks away. An engineer who misses the Kubernetes questions is looking at a longer road, and is better off knowing that now.

How Much Kubernetes Does CKA-Level Knowledge Mean?

CKA-level knowledge means you can administer a Kubernetes cluster from the command line without guidance. Nutanix asks NCP-AI candidates for that level of knowledge, not for the certificate. In practice you should be able to inspect pods and nodes, reason about scheduling, and trace a failed workload to its cause, because the NCP-AI troubleshooting objectives assume all three.

The benchmark itself is public. The CNCF CKA programme is a two hour, performance-based exam solved at a command line, and its largest domain is Troubleshooting at 30 percent. That is a useful clue. The skill Nutanix is borrowing from CKA is mostly diagnosis.

CKA domain CKA weight Where NCP-AI leans on it
Troubleshooting 30% Health check failures, endpoints that will not schedule, image download failures
Cluster Architecture, Installation and Configuration 25% Installation prerequisites and version compatibility
Services and Networking 20% FQDN, ingress and certificate problems after installation
Workloads and Scheduling 15% GPU nodes, taints and allocatable resources
Storage 10% Storage classes and CSI driver connectivity

The scheduling detail that catches people

One NCP-AI objective asks you to determine which allocatable resources, including CPU, memory, GPUs and taints, could stop an endpoint from being scheduled. If the phrase taints and tolerations is new to you, that objective is closed. GPU nodes are commonly tainted so that ordinary workloads stay off them, and an endpoint without the matching toleration simply never starts.

You do not have to hold CKA to pass. However, if you could not pass CKA today, plan to study Kubernetes first and Nutanix Enterprise AI second.

What You Do Not Need: Data Science and Model Training

NCP-AI does not require data science, mathematics or model training skills. The objectives cover importing existing large language models, exposing them through endpoints, and improving output with guardrails and rerank models. You choose and operate models. You never build one, so an infrastructure engineer loses nothing by lacking a machine learning background.

The AI knowledge you do need is practical and narrow. It fits on a short list.

  • Where models come from: the objectives name HuggingFace and NVIDIA NGC, plus a manual import route.
  • What access a model needs: repository keys, and an accepted licence for Llama models.
  • How a model is consumed: an OpenAI-compatible API, called with Python or Curl.
  • How quality is judged: comparing prompt input with output through human feedback.
  • How quality is improved: a different model, guardrails for safety, or a rerank model.

Licences deserve a note, since they produce a named failure in the troubleshooting section. Some models sit behind gated model access, where you must accept terms before a download is allowed. A valid token with an unaccepted licence still fails, and the exam expects you to recognise that.

“AI initiatives are employed to deliver strategic advantages, but those advantages can’t happen without optimized infrastructure control and security.”

Scott Sinclair, Practice Director, ESG

That sentence is the reason the credential exists. Somebody has to own the infrastructure under the model, and this exam checks that they can.

What Do You Commit To When You Book NCP-AI?

Booking NCP-AI commits you to 75 multiple choice questions in 120 minutes at 200 USD per attempt. The passing score is 3000 on a scale of 1000 to 6000. The exam is offered in English and Japanese, and Nutanix announced it could be taken remotely or in person at a PSI testing center.

Field Value
Exam name Nutanix Certified Professional – Artificial Intelligence
Exam code NCP-AI, version 6.10
Questions 75 multiple choice
Duration 120 minutes
Passing score 3000 on a scale of 1000 to 6000
Price 200 USD per attempt
Languages English and Japanese
Recommended course Nutanix Enterprise AI Administration (NAIA)
Sections 5, with no published weightings

Two things follow from those numbers. First, 120 minutes across 75 questions is 96 seconds each, which is enough for recall items and tight for sizing scenarios. Second, a scaled score cannot be converted into a count of correct answers, so you cannot work out a safety margin in advance.

Because no section is weighted, you also cannot decide to skip one. A candidate strong on deployment and weak on troubleshooting has no way of knowing how much that weakness costs. For a decision about readiness, that argues for closing every gap and not gambling on a favourable mix.

NCP-AI, NCP-CN or NCP-MCI: Which Professional Exam Fits?

NCP-AI, NCP-CN and NCP-MCI are all Nutanix professional level exams with the same shape: 75 questions, 120 minutes, 200 USD and a passing score of 3000 on a 1000 to 6000 scale. They differ in subject. Choose NCP-MCI for core infrastructure, NCP-CN for the cloud native track, and NCP-AI for running Nutanix Enterprise AI.

Exam Full name Best first choice if
NCP-MCI Nutanix Certified Professional – Multicloud Infrastructure You administer Nutanix clusters and have little Kubernetes experience
NCP-CN Nutanix Certified Professional – Cloud Native You are moving into Kubernetes on Nutanix and want that base certified
NCP-AI Nutanix Certified Professional – Artificial Intelligence You already have both bases and will operate model endpoints

Since the three exams cost the same and run to the same length, price and effort on the day do not separate them. Your current job does. NCP-AI builds on skills that the other two exams certify directly, so it rewards candidates who arrive with both.

That does not make the other two exams a formal requirement. It makes them a sensible order for someone starting from virtualisation. The site’s Nutanix certification resources list the wider set of Nutanix exams if you want to plan more than one step ahead.

How Do You Close the Gaps in the Right Order?

Close NCP-AI gaps from the bottom of the stack upwards: Kubernetes first, GPUs and storage second, Nutanix Enterprise AI third, and application integration last. Each layer explains the failures of the layer above it, so studying the product before the platform leaves you memorising symptoms you cannot diagnose.

Bottom up study order for NCP-AI: cluster skills first, then GPU and storage sizing, then NAI endpoints
  1. Sort your sample question misses into infrastructure, Kubernetes and product piles, so you know which gap is largest.
  2. Practise Kubernetes at the command line until you can read pod status, node resources and taints without looking anything up.
  3. Learn how GPU nodes are labelled and tainted, and how to confirm whether a workload is really using a GPU.
  4. Install Nutanix Enterprise AI in a lab, including the FQDN and certificate steps, and note every prerequisite you had to meet.
  5. Import one model from a repository with a key, then create an endpoint and size it for a stated throughput.
  6. Create an API key, share the endpoint details, and call the endpoint with Curl and then Python.
  7. Break the lab on purpose with an invalid token, an unaccepted licence and a missing GPU, and trace each fault to its layer.

The recommended course, Nutanix Enterprise AI Administration, covers steps four to six. It will not teach steps two and three from nothing, which is why the order matters.

Book when the sample questions feel like descriptions of things you have already done. At that point the exam is asking you to recall your own work.

Frequently Asked Questions

Is NCP-AI meant for infrastructure engineers or data scientists?

Infrastructure engineers. The exam measures installing, configuring, optimising and troubleshooting Nutanix Enterprise AI and connecting applications to it. No objective covers model training, so a data science background helps far less than Kubernetes and virtualisation experience.

What experience does Nutanix expect for NCP-AI?

At least three years of virtual infrastructure experience and one year with cloud native technologies and the Linux command line. Nutanix also expects knowledge of NCI, cloud IaaS, GPUs and Nutanix Unified Storage.

Do you need the CKA certificate before NCP-AI?

No. Nutanix asks for a Certified Kubernetes Administrator level of knowledge, not the certificate itself. If you could not pass CKA today, study Kubernetes before you study the product.

Do you need NCP-MCI or NCP-CN first?

The official NCP-AI page describes a candidate profile and lists no required prior certification. NCP-MCI and NCP-CN certify skills that NCP-AI assumes, so they make a sensible order for someone starting from virtualisation.

How many questions are on the NCP-AI exam?

The exam has 75 multiple choice questions in 120 minutes, which is 96 seconds per question on average.

What is the NCP-AI passing score?

The passing score is 3000 on a scale of 1000 to 6000. It is a scaled score, so it does not convert into a number of correct answers.

How much does NCP-AI cost?

The price is 200 USD per attempt. NCP-CN and NCP-MCI are listed at the same price.

Which languages is NCP-AI offered in?

English and Japanese. The exam blueprint guide is published in both languages as well.

Does NCP-AI include coding?

Only lightly. One objective asks you to issue a simple query to an OpenAI-compatible endpoint using Python or Curl, and another asks you to tell tool calling commands from non tool calling ones.

How long is the NCP-AI certification valid?

Neither the syllabus page nor the official exam page used for this article states a validity period, so confirm the current renewal rule with Nutanix University before you plan around it.

Conclusion

NCP-AI is an infrastructure credential, and the decision about sitting it is a decision about your infrastructure base. Three years of virtualisation, a year of cloud native work and administrator-level Kubernetes are the conditions Nutanix names, and the objectives show that each one is there for a reason.

If you meet them, the remaining work is the product: models, endpoints, keys and metrics across 75 questions in 120 minutes. If you do not, the honest move is to close the Kubernetes gap first and come back. Either way, test yourself on real questions before paying 200 USD.

For engineers who already hold the professional tier and want to see what the next level demands, the guide to the NCM-MCI master exam shows where the Nutanix ladder goes after this.

Rating: 0 / 5 (0 votes)

The post NCP-AI Prerequisites: Is the Nutanix AI Exam for You? appeared first on Big Data Prep.

]]>
Tableau Architect Exam: 8 Sign-In Methods, 7 Migrations https://www.bigdataprep.com/2026/09/28/tableau-architect-exam-8-sign-in-methods-7-migrations/ Mon, 28 Sep 2026 00:00:00 +0000 https://www.bigdataprep.com/?p=16429 Servers, identity providers and moving deployments between environments, named task by named task: that is the Analytics-Arch-201 objective list.

The post Tableau Architect Exam: 8 Sign-In Methods, 7 Migrations appeared first on Big Data Prep.

]]>

The Salesforce Certified Tableau Architect exam, Analytics-Arch-201, asks one question over and over: can you run Tableau as a platform? Sign-in methods, node roles, migrations between environments and the monitoring that keeps a deployment healthy fill the paper, so a strong workbook author with no server experience can sit it and recognise almost nothing.

What the Tableau architect exam covers instead is the platform underneath the workbooks. Its objective list names eight authentication methods to configure and troubleshoot, seven separate migrations to plan, a set of external services to stand up, and a load testing routine to run. It is 59 questions in 105 minutes at a 63 percent pass mark, and 78 percent of it sits in deploying, monitoring and maintaining a Tableau deployment rather than designing one.

Table of Contents

  1. Architect, administrator or consultant: which Tableau exam is this?
  2. What are the Analytics-Arch-201 exam facts?
  3. How are the three domains weighted?
  4. What does the 22 percent design domain ask you to plan?
  5. Which eight authentication methods must you configure?
  6. Why is monitoring and maintenance worth 41 percent?
  7. What does Salesforce expect before you book?
  8. How should you prepare for the Tableau architect exam?
  9. Frequently Asked Questions
  10. Conclusion

Architect, administrator or consultant: which Tableau exam is this?

Analytics-Arch-201 is the Tableau exam for people who build and scale the platform itself. Salesforce aims it at technical architects who lead the design of a Tableau Server deployment or a Tableau Cloud migration. The consultant exam covers advising on analytics content, and the server administrator exam covers running an existing site.

The three are easy to confuse because their titles all suggest seniority. The practical test is what your working week looks like. If you spend it deciding node roles, wiring identity providers, sizing backgrounder processes and planning how a deployment moves from one environment to another, the architect exam describes you. If you spend it recommending dashboards, calculations and governance to business teams, the consultant route fits better.

The administrator credential sits closer to the architect one, and that is where the real decision lies. Day to day administration of users, schedules and sites is assumed knowledge here rather than the thing being tested. The architect paper asks what you would build, how you would move it, and how you would prove it is healthy once it is running.

  • Choose the architect exam if you design multi node Tableau Server deployments or lead Tableau Cloud migrations
  • Choose the administrator exam if you run a Tableau Server or Tableau Cloud site that somebody else designed
  • Choose the consultant exam if your deliverable is analytics content and recommendations for business users

What are the Analytics-Arch-201 exam facts?

Analytics-Arch-201 has 59 questions in 105 minutes and a 63 percent passing score. Registration costs 400 US dollars and a retake costs 200. Of the 59 questions, 54 are scored multiple choice and up to five are unscored, and the exam tests Tableau product version 2024.2.

Field Value
Credential Salesforce Certified Tableau Architect
Exam code Analytics-Arch-201
Questions 59 (54 scored, up to 5 unscored)
Duration 105 minutes
Passing score 63 percent
Registration fee 400 USD plus applicable taxes
Retake fee 200 USD plus applicable taxes
Delivery Proctored, at a test centre or online, registered through Pearson VUE
Prerequisites None
Reference materials None permitted during the exam
Product version Tableau 2024.2
Maintenance Annual maintenance module on Trailhead

The unscored questions matter more than they look. You cannot tell which five they are, so every question has to be treated as live, and the pass mark applies to the 54 that count. In practice that means pacing for all 59 inside the 105 minutes, which leaves a little under two minutes a question.

The product version is the second detail worth noting. Salesforce’s official exam guide states that the exam currently tests release 2024.2, so features added to Tableau Server and Tableau Cloud after that release are not what the questions are written against.

How are the three domains weighted?

The Analytics-Arch-201 exam has three domains. Design a Tableau Infrastructure is 22 percent, Deploy Tableau Server is 37 percent, and Monitor and Maintain a Tableau Deployment is 41 percent. Deployment and ongoing operation together account for 78 percent of the exam, so design is the smallest domain despite the architect title.

Domain Weight Objective groups Headline content
Design a Tableau Infrastructure 22% 5 Requirements, Tableau Cloud, migrations, topology, configuration
Deploy Tableau Server 37% 5 Production deployments, authentication, encryption, Linux and Windows installs
Monitor and Maintain a Tableau Deployment 41% 6 Admin views, load testing, performance, observability, automation, extensions

Salesforce weights by domain, not by objective group, so the table tells you where effort belongs rather than how many questions any single topic will produce. The money site and Salesforce publish the same three domains, the same weightings and the same sixteen objective groups, and the full Arch-201 syllabus lists every named task under each group.

Reading that list end to end is the fastest way to see what kind of exam this is. The verbs are configure, troubleshoot, install, implement, interpret and automate. Recommend appears too, but it is attached to concrete things: configuration keys, hardware specifications, a disaster recovery strategy, a load testing strategy.

What does the 22 percent design domain ask you to plan?

The design domain in the Tableau architect exam covers five areas: gathering requirements for a complex deployment, planning and implementing Tableau Cloud, planning migrations, designing a process topology, and recommending a Tableau Server configuration. Migration planning alone names seven separate scenarios, which makes it the most specific part of the domain.

Tableau architect exam migration directions: Cloud to Server, Server to Cloud, Windows to Linux and Linux to Windows

Requirements and Tableau Cloud

Requirements gathering here is infrastructure requirements, not business ones. You assess user counts and role distribution, constraints and future growth, the need for high availability and disaster recovery, and the licensing strategy, including Authorization-to-Run. You also map the Tableau Server add-ons to what the organisation actually needs.

The Tableau Cloud objective adds Tableau Bridge, authentication, multi-site management through Tableau Cloud Manager, and automated user provisioning. Provisioning is named with its standard, System for Cross-domain Identity Management, whose SCIM protocol specification is published by the IETF. Knowing what SCIM does and does not synchronise is the difference between a working provisioning design and a support queue.

The seven migrations

  • Tableau Cloud to Tableau Server
  • Tableau Server to Tableau Cloud
  • Windows to Linux
  • Linux to Windows
  • One identity store to another
  • Several Tableau servers or sites consolidated into fewer
  • One Tableau Server environment to another

The objective adds two tools to that list: writing scripts for migration, and using the Tableau Content Migration Tool. Tableau’s own Content Migration Tool documentation carries a detail worth knowing: the tool is part of Tableau Advanced Management, and Tableau does not recommend it for moving from Tableau Server to Tableau Cloud, pointing instead to the Tableau Migration SDK. Choosing the right tool for the direction of travel is exactly the kind of judgement this objective tests.

Topology and configuration

The last two objective groups are where sizing happens. You specify process counts, node count and node roles, including when to isolate a service and when to colocate it, and when to move a component to an external service. Configuration recommendations cover the identity store, specific configuration keys, encryption at rest and over the wire, hardware and network specifications, and a disaster recovery strategy.

Which eight authentication methods must you configure?

The Tableau architect exam names eight authentication methods to configure and troubleshoot: SAML, Kerberos, OpenID Connect, Mutual SSL, trusted authentication, Connected App authentication, LDAP, and Azure Active Directory. They sit inside Deploy Tableau Server, the 37 percent domain, which also covers production deployments, encryption, and installation on both Linux and Windows.

Authentication is a troubleshooting objective

Every method in that list carries the same pair of verbs: configure and troubleshoot. Knowing that Tableau supports SAML is not enough; the exam expects you to know why a SAML sign-in fails and where to look. It helps to understand the protocol itself, and the OASIS SAML standard is the specification of record for the assertions an identity provider sends.

The objective closes with dependencies between authentication methods and Tableau environments, including Tableau Cloud. Not every method is available everywhere, and some cannot be combined, so a question may describe an environment and ask which method will actually work in it.

Production deployments and encryption

The production deployment objective is a long list of named components. You are expected to configure an external file store, an external repository, an external gateway, an unlicensed node, a coordination ensemble, a backgrounder with a specific node role, and a load balancer. Beyond configuration, it names installing in an air-gapped environment, performing a blue-green deployment, validating a disaster recovery and high availability test, reading installation logs, running the Resource Monitoring Tool, and automating installation with the Silent Installer.

Encryption is its own objective group: SSL, database encryption, extract encryption, and setting up service principal names for Kerberos. Tableau’s guide to distributed and high availability installations is the reference for how these pieces fit together, and it is candid that Tableau Cloud will suit most organisations better than self-hosting, which is why Cloud migration keeps appearing in an exam about servers.

Linux and Windows, both

Installation is tested twice, once per operating system, with near-identical objectives: install by command line or wizard, resolve installation, network, external system and proxy issues, find the right logs, and verify system groups and file permissions. Windows adds the Run As service account. If your experience is on one platform only, the other is an obvious gap to close.

Why is monitoring and maintenance worth 41 percent?

Monitor and Maintain a Tableau Deployment is the largest Analytics-Arch-201 domain at 41 percent because it covers six objective groups: custom administrative views, load testing, performance bottlenecks, observability data, automation of server maintenance, and server extensions. It tests whether you can keep a deployment healthy and prove it with evidence, not just build it.

Evidence first: admin views, load tests and observability

Custom administrative views start from the repository schema and its event types, and extend to building admin dashboards and using Admin Insights on Tableau Cloud. This is the one place the exam asks you to build a dashboard, and it is a dashboard about the server rather than the business.

Load testing is named in detail: a strategy, a tool such as TabJolt, a test environment, test plans, and interpreting results to decide what to do next. Observability goes further, asking you to collect and analyse logs, process metrics and operating system metrics, interpret them, and revise the architecture based on what they show. That last step joins this domain back to design.

Performance and automation

  • Performance: diagnose slow workbooks and data sources, run resource, latency and workload analysis, act on performance recordings, and optimise caching
  • Automation: manage the server through TSM, REST APIs and tabcmd, schedule scripts with Windows Scheduler or cron, and automate backup and cleanup
  • Upgrades and recovery: plan multi node upgrades and design an automated disaster recovery process
  • Metadata: configure and use the Metadata API
  • Extensions: schedule content automation with webhooks, tabcmd, REST or Hyper APIs, enable dashboard extensions and web data connectors, and configure trusted tickets and connected apps for embedding

The common thread is that nothing here is a one-off task. A candidate who has installed Tableau Server once but never operated it through a quarter of growth, an upgrade and a performance incident will find this domain the hardest, and it is the one that carries the most marks.

What does Salesforce expect before you book?

Salesforce recommends at least two years of Tableau administration across Tableau Cloud, Tableau Server and Tableau Bridge, plus having architected Tableau Server in at least one distributed environment, whether public cloud, private cloud or on-premises. There are no formal prerequisites, but the objectives assume that experience throughout.

Tableau architect exam checklist before booking: two years of Tableau administration, a multi node build, product version 2024.2 and yearly maintenance

The distributed environment requirement is the one to check honestly. A single node installation never forces decisions about node roles, a coordination ensemble, an external repository or a load balancer, and those decisions are scattered across two of the three domains. If every Tableau Server you have touched ran on one machine, build a multi node lab before booking.

The credential also comes with a running commitment. Holders complete a Tableau Architect maintenance module on Trailhead once a year to keep the certification current, so passing is the start of a yearly cycle rather than a one-time event.

The site has covered this credential before under its older code, and the TCA-C01 architect companion is useful for the role framing, though its figures predate the current exam and should not be used for planning.

How should you prepare for the Tableau architect exam?

Prepare for Analytics-Arch-201 in a lab rather than a reader. Build a multi node Tableau Server, wire up more than one authentication method, run a migration and a load test, then study in proportion to the weightings, with roughly four fifths of your time on deployment, monitoring and maintenance.

  1. Map your experience against the sixteen objective groups and mark every named task you have never performed yourself.
  2. Build a multi node Tableau Server lab with an external repository, an external file store and a coordination ensemble.
  3. Configure at least three of the eight authentication methods, then break each one on purpose and fix it from the logs.
  4. Plan one migration in each direction between Tableau Server and Tableau Cloud, and note which tool fits each direction.
  5. Run a load test against the lab, interpret the results, and change the topology in response.
  6. Automate a backup, a cleanup and a scheduled content task through TSM, tabcmd and the REST API.
  7. Finish with timed practice of 59 questions in 105 minutes to check your pacing before the real sitting.

Two groups of candidates need to adjust the plan. Administrators who have only run Tableau Cloud should spend extra time on Windows and Linux installation, encryption and topology, since none of that exists on a hosted site. People whose Tableau work is analytics advice rather than platform work should reconsider the target altogether, because the Tableau Consultant guide describes an exam built around recommendations on content, which may fit that background far better.

Frequently Asked Questions

How many questions are on the Analytics-Arch-201 exam?

The exam has 59 questions. Salesforce states that 54 are scored multiple choice questions and up to five are unscored, and the unscored ones do not affect your result.

What is the passing score for the Tableau architect exam?

The passing score is 63 percent, and you have 105 minutes to complete the exam.

How much does the Salesforce Certified Tableau Architect exam cost?

Registration costs 400 US dollars and a retake costs 200 US dollars, both plus applicable taxes. The exam is registered through Pearson VUE.

Which domain carries the most weight?

Monitor and Maintain a Tableau Deployment at 41 percent. Deploy Tableau Server is 37 percent and Design a Tableau Infrastructure is 22 percent.

Are there prerequisites for Analytics-Arch-201?

There are no formal prerequisites. Salesforce recommends at least two years of Tableau administration and experience architecting Tableau Server in at least one distributed environment.

Which Tableau version does the exam test?

The official exam guide states that the exam currently tests product version 2024.2.

Does the exam cover Tableau Cloud or only Tableau Server?

Both. Tableau Cloud appears in its own design objective, in the migration plans, in authentication dependencies and in Admin Insights, although most objectives concern Tableau Server.

How do you keep the certification current?

Holders complete the Tableau Architect maintenance module on Trailhead once a year.

Can you use notes or documentation during the exam?

No. Salesforce states that no hard copy or online materials may be referenced while you sit the exam.

Conclusion

The Tableau architect exam is a platform exam. Fifty nine questions in 105 minutes at 63 percent, 400 US dollars to register, and three domains weighted 22, 37 and 41 percent, with deployment and operations taking 78 percent between them.

Its most useful feature is how specific it is. Eight named authentication methods, seven named migration plans, a named list of external services, and a load testing and observability routine give you a checklist you can test yourself against before spending the fee. Anything on that list you have only read about is the place to start.

Build the multi node lab, break the sign-in methods on purpose, move content in both directions between Server and Cloud, and measure the deployment before you tune it. That work prepares you for the questions and for the job the credential describes.

Rating: 0 / 5 (0 votes)

The post Tableau Architect Exam: 8 Sign-In Methods, 7 Migrations appeared first on Big Data Prep.

]]>
Snowflake Data Scientist Certification: Python Runs Inside https://www.bigdataprep.com/2026/09/23/snowflake-data-scientist-certification-python-runs-inside/ Wed, 23 Sep 2026 00:00:00 +0000 https://www.bigdataprep.com/?p=16413 Vector embedding, fine tuning and task-specific language models are examinable topics on this credential, sitting in its heaviest domain. Most study material still describes the syllabus that came before them.

The post Snowflake Data Scientist Certification: Python Runs Inside appeared first on Big Data Prep.

]]>

Prompt engineering is examinable on a data scientist credential now. So is vector embedding, so is fine tuning, and they sit inside the largest domain of the exam rather than in an appendix. Most summaries of this certification still describe a syllabus that stopped at model deployment, and they are describing the previous version.

The Snowflake data scientist certification, exam code DSA-C03, is 65 questions in 115 minutes at 375 dollars, scored on a 0 to 1000 scale with 750 to pass. Its four domains run from data science concepts at 17 percent through feature engineering and model development to deployment, and generative AI now sits inside the largest of them.

Table of Contents

  1. What does the Snowflake data scientist certification examine?
  2. Why is prompt engineering on a data science exam?
  3. What does the DSA-C03 exam look like?
  4. What do the four DSA-C03 domains cover?
  5. Why does model development carry the largest share?
  6. Feature engineering happens in Snowpark, not in a notebook
  7. What does deployment mean when the model stays in the warehouse?
  8. Who is DSA-C03 for, and how should you prepare?
  9. Frequently Asked Questions
  10. Conclusion

What does the Snowflake data scientist certification examine?

DSA-C03 examines whether you can do data science inside Snowflake rather than alongside it. Snowflake states that the credential validates five capabilities: outlining data science concepts, implementing Snowflake data science best practices, preparing data and using feature engineering, training and using machine learning models, and using generative AI and large language model capabilities.

That last capability is the newest and the most consequential. The other four describe a recognisable data science workflow; the fifth describes something that did not belong on this kind of exam two years ago.

The word doing the heavy lifting throughout is “in Snowflake”. This is not a general machine learning exam that happens to mention a warehouse. The objectives name Snowpark, Snowpark ML, Python user-defined functions, stored procedures, user-defined table functions, dynamic tables, the Snowpark Feature Store, the Model Registry and Snowpark Container Services. The platform is the subject, not the setting.

The money site’s DSA-C03 certification page sets the four weighted domains out beside the exam terms, which is the quickest way to see how much of the syllabus is platform-specific.

Why is prompt engineering on a data science exam?

Because Snowflake has put large language models inside the warehouse, and the data scientist is the person expected to use them. The model development domain names Snowflake Cortex, vector embedding, prompt engineering, fine tuning, and task-specific models for categorisation, summarisation, sentiment and information extraction.

The DSA-C03 model development domain holds classic machine learning tuning and validation alongside the new Cortex embedding, prompting and fine tuning content

Read that list next to the traditional content in the same domain, which covers hyperparameter tuning, cross validation and optimisation metrics, and the intent becomes clear. Snowflake is treating a call to a hosted language model as one more technique a data scientist reaches for, sitting beside a gradient-boosted tree rather than in a separate discipline.

What this means for preparation

A candidate who learned Snowflake data science before Cortex existed has a genuine gap, and it is in the heaviest-weighted part of the paper. The Cortex LLM function reference is the fastest way to close it, because the exam names the task-specific functions rather than the theory behind them.

For anyone who wants the generative material in depth rather than as one domain among four, the Gen AI specialty credential covers exactly this ground as its whole syllabus, and the overlap between the two is now substantial.

Vector embedding is the concept to get right

Embedding is the mechanism that lets unstructured text sit next to structured data in the same query, and almost every practical Cortex pattern depends on it. It is one line in the objectives and it underpins retrieval, similarity and classification workflows alike, which makes it worth more attention than its single mention suggests.

What does the DSA-C03 exam look like?

DSA-C03 is 65 questions in 115 minutes, delivered through Pearson VUE at 375 dollars per attempt, and scored on a scale from 0 to 1000 with 750 required to pass. That works out at roughly 106 seconds a question, which is tight for a paper this technical.

Field Value
Exam name SnowPro Advanced: Data Scientist
Exam code DSA-C03
Questions 65
Duration 115 minutes
Passing score 750, on a scale of 0 to 1000
Price USD 375 per attempt
Delivered by Pearson VUE
Candidate profile 2 or more years of hands-on Snowflake experience as a data scientist in production
Domains 4, weighted

A scaled score of 750 out of 1000 is not the same as 75 percent of the questions. Scaling exists so that different forms of the exam are measured against one standard even though the question sets differ, and Snowflake does not publish the transformation, so there is no way to convert 750 back into a number of correct answers.

Two things this article does not tell you

Snowflake’s certification page renders its overview and its pricing, but its exam details panel does not load for anything other than a browser. That means the validity period and any prerequisite credential could not be verified, and neither is asserted here. Both are worth confirming on the official DSA-C03 page before you register.

What do the four DSA-C03 domains cover?

DSA-C03 has four weighted domains: data science concepts at 17 percent, data preparation and feature engineering at 27 percent, model development at 31 percent and model deployment at 25 percent. The two middle domains together carry 58 percent, and both are heavily Snowflake-specific.

The four DSA-C03 stages a model passes through: prepared in Snowpark, built in the warehouse, validated on scores and curves, served through the registry
Domain Weight What it examines
Data Science Concepts 17% Supervised and unsupervised learning, problem types from regression to time-series forecasting and image segmentation, the machine learning lifecycle, evaluation measures including the confusion matrix, and statistical foundations including distributions, outliers and the central limit theorem
Data Preparation and Feature Engineering 27% Cleaning and preparing data with Snowpark and SQL, exploratory analysis and profiling, Snowflake’s native statistical functions, preprocessing through scaling and encoding, DataFrames across Pandas and Snowpark, the Snowpark Feature Store, and presenting findings through Snowsight and Snowflake Notebooks
Model Development 31% Connecting Python to Snowflake, Cortex and the generative AI stack, building pipelines with dynamic tables and Python functions, hyperparameter tuning, optimisation metrics, cross validation and sampling, validation with ROC curves and residuals, and interpretation through feature impact and partial dependence
Model Deployment 25% External hosted models and external functions, in-Snowflake deployment with vectorised and scalar Python functions, the Model Registry, Snowpark Container Services, data drift and model decay, retraining, metadata tagging and versioning

The smallest domain is the one most candidates are already strongest in. Data science concepts at 17 percent is standard theory that transfers from any other machine learning background, which means roughly 83 percent of the exam depends on knowing Snowflake specifically.

Why does model development carry the largest share?

Model development is 31 percent because it holds three separate bodies of work rather than one. It covers getting Python connected to Snowflake at all, the entire generative AI stack, and the traditional business of training, validating and interpreting a model. Any one of those would sustain a domain on its own.

The connection objectives alone name Snowpark, Snowpark ML, the Python connector with Pandas support, the Spark connector and connecting from an external development environment. Knowing which of those is appropriate to which situation is a recurring question shape, and it depends on understanding where the computation actually runs.

Training has several homes

The objectives name training with Python stored procedures, training with user-defined table functions, and training outside Snowflake through external functions. Those are three genuinely different architectures with different cost, scaling and governance consequences, and the exam expects you to be able to choose between them rather than to know one.

Validation and interpretation are examined separately

Validation covers the ROC curve, the confusion matrix, expected payout, residuals plots and model metrics. Interpretation is its own objective and covers feature impact, partial dependence plots and confidence intervals. Grouping them in your head is a mistake, because one asks whether the model works and the other asks why it produces the answers it does.

Optimisation metric selection is worth singling out, since the objectives name log loss, area under the curve and root mean squared error explicitly. Choosing the wrong metric for a problem type is a classic question, and the three named here map onto probabilistic classification, ranking and regression respectively.

Feature engineering happens in Snowpark, not in a notebook

Data preparation and feature engineering carries 27 percent, and the defining feature of the domain is where the work happens. Every preparation objective is framed around Snowpark for Python and SQL, native statistical functions and Snowflake Notebooks rather than around pulling an extract onto a laptop.

The practical consequence is that a data scientist who is fluent in the standard Python stack but has always exported data first will find the domain harder than the topic names suggest. Scaling, encoding, normalisation, binning and one-hot encoding are all familiar; doing them against a DataFrame that is actually executing as SQL is not.

Three DataFrame flavours, and the exam names all three

The objectives list Pandas, Snowpark and Snowpark pandas as separate things. That distinction is the single most useful thing to internalise in this domain: a pandas DataFrame holds data in memory on the client, a Snowpark DataFrame is a lazily-evaluated query plan that runs in the warehouse, and Snowpark pandas offers the familiar interface over the second of those. The Snowpark developer guide is the reference that makes the difference concrete.

Native statistical functions are free marks

Window functions, MIN, MAX, AVG, STDEV, VARIANCE, TOPn and the approximation functions are all named in the objectives, and they are the most learnable content in the domain. Anyone comfortable with SQL can close this gap in an evening, which makes it a good early win in a study plan.

The Feature Store is the newer item here. It is named once, and it is the mechanism for keeping engineered features consistent between training and serving, which is a problem most practitioners have solved badly by hand at some point.

What does deployment mean when the model stays in the warehouse?

Model deployment carries 25 percent and it is the domain that looks least like a general machine learning exam. The objectives split between putting a model into production and keeping it honest afterwards, and both halves are expressed in Snowflake mechanisms.

Production deployment covers external hosted models and external functions on one side, and vectorised and scalar Python user-defined functions, stored predictions, stage commands, the Model Registry and Snowpark Container Services on the other. The vectorised against scalar distinction matters in practice, because one processes a batch per call and the other a row, with very different throughput.

Model decay is examined as a first-class topic

Data drift and model decay appear by name, along with data distribution comparisons framed as two questions: does the data making predictions still look like the training data, and do the same inputs still produce the same outputs after deployment. That framing is unusually concrete for a syllabus, and it tells you the exam wants operational thinking rather than a definition.

Versioning and retraining close the loop

Metadata tagging, model versioning in the Model Registry and automation of retraining are the final objectives, and together they describe the lifecycle the first domain introduced in the abstract. Candidates who study the four domains as separate subjects miss that the last objective of the exam answers the first.

Who is DSA-C03 for, and how should you prepare?

Snowflake states the expected candidate has two or more years of hands-on Snowflake experience as a data scientist in a production environment, and may have worked in Python, R, SQL or PySpark. That is a genuine bar rather than a suggestion, and roughly 83 percent of the exam depends on platform knowledge that only production work builds.

  1. Confirm you actually meet the candidate profile before spending 375 dollars, because the weighting means general machine learning strength will not carry a thin Snowflake background.
  2. Start with the Snowflake-native statistical functions and window functions, which are the most learnable content in the heaviest preparation domain and give an early sense of progress.
  3. Rebuild one familiar feature engineering pipeline entirely in Snowpark, so that scaling, encoding and binning stop being notebook habits and become warehouse operations.
  4. Work through the three DataFrame types deliberately, comparing Pandas, Snowpark and Snowpark pandas on the same task and noticing where the computation happens in each.
  5. Train one model three ways, through a Python stored procedure, through a user-defined table function, and through an external function, since the exam asks you to choose between those architectures.
  6. Spend real time in Cortex, covering vector embedding, prompt engineering, fine tuning and at least two task-specific functions, because this is the newest content in the largest domain.
  7. Deploy a model into the Model Registry and serve predictions through both a scalar and a vectorised user-defined function, which makes the throughput distinction obvious.
  8. Set up a drift check by comparing a live prediction distribution against the training distribution, so the monitoring objectives are something you have done rather than read.

If you have not done SnowPro Core

The Advanced line assumes the platform fundamentals that the SnowPro Core credential covers, including warehouses, storage, roles and the query model. Whether Core is formally required could not be verified from any readable source, so confirm that with Snowflake, but the knowledge is assumed either way.

Frequently Asked Questions

How many questions are on the DSA-C03 exam?

Sixty-five questions in 115 minutes, which works out at roughly 106 seconds each.

What is the passing score for DSA-C03?

Seven hundred and fifty on a scale of 0 to 1000. Because the scaling transformation is not published, that does not translate into a fixed number of correct answers.

What does DSA-C03 cost?

USD 375 per attempt, which is the price across the SnowPro Advanced series. SnowPro Core sits at USD 175.

Does DSA-C03 cover generative AI?

Yes, and it is in the largest domain. Snowflake Cortex, vector embedding, prompt engineering, fine tuning and task-specific models all sit inside model development.

What are the DSA-C03 domains and weightings?

Data science concepts 17 percent, data preparation and feature engineering 27 percent, model development 31 percent and model deployment 25 percent.

What experience does Snowflake expect?

Two or more years of hands-on Snowflake experience as a data scientist in a production environment, with possible experience in Python, R, SQL or PySpark.

Is SnowPro Core required before DSA-C03?

Snowflake’s exam details panel does not render to anything but a browser, so no prerequisite is asserted here. Confirm it on the official page. The platform knowledge Core covers is assumed by the syllabus regardless.

How long is the certification valid?

No validity period is visible on any readable source, so none is stated here. Check Snowflake’s certification page before assuming a term.

Is DSA-C03 replacing DSA-C02?

DSA-C03 is the current code and the one Snowflake’s certification page names. Third-party material describing the credential under the older code is describing the previous syllabus, which did not include the generative AI content.

How long does preparation usually take?

Eight to twelve weeks alongside a job for somebody already working in Snowflake, with the largest single block going to Cortex if the generative content is new to you.

Conclusion

DSA-C03 is a Snowflake exam before it is a data science exam. Only 17 percent of it is portable theory; the rest depends on knowing how Snowpark, user-defined functions, stored procedures, the Feature Store, the Model Registry and Container Services actually behave.

The change worth acting on is the generative AI content in model development. Prompt engineering, vector embedding and fine tuning are examinable, they sit in the heaviest-weighted domain, and most study material still describes a syllabus that predates them.

Prepare by moving work you already do into Snowpark rather than by reading about it, spend a deliberate block on Cortex, and treat the 750 scaled mark as what it is: a standard, not a percentage you can compute your way to.

Rating: 0 / 5 (0 votes)

The post Snowflake Data Scientist Certification: Python Runs Inside appeared first on Big Data Prep.

]]>
SAS Regression and Modeling Certification: Beyond the Fit https://www.bigdataprep.com/2026/09/16/sas-regression-and-modeling-certification-beyond-the-fit/ Wed, 16 Sep 2026 00:00:00 +0000 https://www.bigdataprep.com/?p=16402 Fitting a model accounts for just over half this paper. The rest is the work either side of the fit, and those are the domains an analyst is least likely to have picked up informally. Here is how the five domains divide, and why input preparation and model measurement should be studied as one subject.

The post SAS Regression and Modeling Certification: Beyond the Fit appeared first on Big Data Prep.

]]>

The name of this exam describes about half of what is in it. Regression and modeling is right there in the title, and yet 45 percent of the marks fall after the model has been fitted, in the work of preparing what goes into it and judging whether what came out is any good.

The SAS regression and modeling certification, exam code A00-240, is 60 questions in 110 minutes at a 68 percent pass mark, and its five weighted domains split almost evenly between building models and assessing them.

Table of Contents

  1. What is the SAS regression and modeling certification?
  2. Why does the exam name describe only half of it?
  3. How is A00-240 delivered and scored?
  4. Where do the marks sit across the five domains?
  5. Why is ANOVA mostly about assumptions?
  6. Linear and logistic regression together are 45 percent
  7. What does preparing inputs actually involve?
  8. What does measuring model performance involve?
  9. How should you prepare for A00-240?
  10. Frequently Asked Questions
  11. Conclusion

What is the SAS regression and modeling certification?

The SAS regression and modeling certification is SAS’s advanced analytics credential for analysts who build and evaluate statistical models in SAS 9, formally the SAS Certified Statistical Business Analyst Using SAS 9: Regression and Modeling and carrying the exam code A00-240. It covers ANOVA, linear regression, logistic regression, input preparation and model assessment across 60 questions.

SAS files it under Advanced Analytics rather than under programming, and lists it among its most popular credentials. That placement is the clearest statement of what the exam is for: it assumes you can already write SAS and asks whether you can do statistics with it.

What distinguishes it from a general statistics qualification is how specific the syllabus is. It does not ask whether you understand analysis of variance in the abstract. It names the procedures, the statements and the options, down to which option of which statement performs a particular test.

Why does the exam name describe only half of it?

Add up the domains that are about fitting a model and you get 55 percent: ANOVA at 10, linear regression at 20 and logistic regression at 25. Add up the two that are not, Prepare Inputs for Predictive Model Performance at 20 and Measure Model Performance at 25, and you get 45 percent.

The A00-240 split between building the model at 55 percent and everything around it at 45 percent

That is a much more even split than the title suggests, and it changes who the exam is hard for. An analyst who fits models daily but never formally assesses them is missing nearly half the paper. An analyst who does the full modelling cycle, including honest assessment, is already most of the way there.

The two non-fitting domains are also the ones least likely to have been learned informally. Fitting a model is what a course teaches. Deciding which candidate inputs belong in it, and then measuring whether the fitted model actually performs, tends to be learned on the job or not at all.

The quickest way to find out which side you are on is to work a mixed set of exam-style items and notice which ones slow you down. The sample sets on the money site’s A00-240 practice exam mix the five domains in roughly their published proportions, which makes the gap visible in an hour rather than on results day.

How is A00-240 delivered and scored?

A00-240 is 60 questions in 110 minutes with a 68 percent pass mark, priced at $180 USD and delivered through Pearson VUE. SAS publishes the $180 figure on its own certification page, and the money site’s syllabus supplies the question count, duration and pass mark that SAS does not publish anywhere.

Field Value
Credential name SAS Certified Statistical Business Analyst Using SAS 9: Regression and Modeling
Exam code A00-240
Questions 60
Duration 110 minutes
Passing score 68 percent
Price $180 USD
Delivery Pearson VUE
Domains 5, all weighted
SAS category Advanced Analytics
Prerequisite certification None published

Sixty questions in 110 minutes gives 110 seconds each, which is generous for a multiple-choice paper and is clearly deliberate. Several objectives ask you to interpret output rather than recall a fact, and reading a table of parameter estimates or a diffogram properly takes longer than answering a definition.

Sixty eight percent of 60 questions means 41 correct answers and a margin of 19. That is a middling allowance, and the even domain spread means it cannot absorb a whole missing domain: the smallest domain alone is worth six questions and the largest fifteen.

SAS keeps its detail thin on the web. Its certification page names the credential and the price and stops, with no per-credential page behind it, which is why the published domain weightings come from the money site.

Where do the marks sit across the five domains?

Logistic Regression and Measure Model Performance tie as the largest domains at 25 percent each, Linear Regression and Prepare Inputs tie at 20 percent each, and ANOVA is smallest at 10 percent. On a 60 question paper that is 15, 15, 12, 12 and 6 questions respectively.

Domain Weight Approximate questions What it covers
Logistic Regression 25% 15 Binary outcome modelling and the procedures that fit it
Measure Model Performance 25% 15 Assessing a fitted model against held-out data
Linear Regression 20% 12 Multiple linear models, fit and diagnostics
Prepare Inputs for Predictive Model Performance 20% 12 Getting candidate variables ready before modelling
ANOVA 10% 6 Assumption checking, group mean comparison and interaction

Two things follow from that shape. The first is that logistic regression outweighs linear regression, which surprises people who assume the simpler technique carries more. The second is that the two assessment domains together outweigh either regression domain individually.

There is no domain small enough to ignore. Even ANOVA at 10 percent is worth six questions against a 19 question margin, so writing it off consumes nearly a third of the allowance before the paper starts.

Why is ANOVA mostly about assumptions?

The smallest domain is also the most conceptually front-loaded. Before it reaches any group comparison, the syllabus asks about the central limit theorem, the distribution of continuous variables through histograms, box-whisker plots and Q-Q plots, the effect of skewness, the null and alternative hypotheses, Type I and Type II error, statistical power, and how sample size affects both p-value and power.

Only after that does it get to the procedures. The objectives name PROC GLM with its CLASS, MODEL, MEANS and OUTPUT statements, PROC TTEST for comparing means, the HOVTEST option of the MEANS statement for assessing equal response variance, and PROC UNIVARIATE for examining residuals.

The post hoc objectives are more specific still: LSMEANS with the PDIFF option for pairwise comparisons, the ADJUST option using TUKEY and DUNNETT, and interpreting diffograms and control plots to evaluate those comparisons. Knowing that Tukey compares every pair while Dunnett compares against a control is exactly the kind of distinction a question can turn on.

The domain closes with interactions, where PROC PLM and the SLICE= option appear alongside Type I and Type III sums of squares. Six questions is not many for that much named machinery, which is why this domain rewards a focused sweep rather than deep study.

Linear and logistic regression together are 45 percent

The two regression domains hold 27 of the 60 questions between them, and they are weighted the way the working world weights them rather than the way a textbook orders them. Logistic gets 25 percent to linear’s 20, because a binary outcome is what most business modelling actually predicts.

Comparison of linear and logistic regression on the A00-240 syllabus and how their output is read

Linear regression

The linear domain is built around fitting multiple models with PROC REG and PROC GLM, and then around everything that follows a fit: which predictors earned their place, whether the residuals behave, and whether the model generalises. It is the domain where an analyst’s informal habits are most likely to be tested against formal practice.

Logistic regression

Logistic carries the extra five points and deserves the extra attention. The recommended SAS course list names Predictive Modeling Using Logistic Regression specifically, which is a strong signal about the depth expected. SAS’s documentation hub is the place to read the procedure detail, since the statistical procedure pages themselves are the reference the objectives are written against.

The conceptual jump candidates underestimate is interpretation. A linear coefficient is a change in the outcome. A logistic coefficient is a change in log odds, and converting that into something a business audience can act on is a separate skill from fitting the model.

What does preparing inputs actually involve?

Prepare Inputs for Predictive Model Performance is 20 percent, about 12 questions, and it covers the work that happens between having data and having a model worth fitting. It is the domain most often skipped in self-study because it feels like preparation rather than technique.

In practice it is technique. Deciding how to handle missing values, how to treat outliers and extreme values, how to represent categorical variables, and which candidate inputs to carry forward are all decisions with consequences that show up much later, in the assessment domain, as a model that fits the training data and nothing else.

That connection is why the two domains sit adjacent on the syllabus and together carry 45 percent. Input preparation is where overfitting is created, and model measurement is where it is discovered. Reading the two as one continuous subject rather than as two separate ones is the most useful reframing available for this exam.

Anyone whose SAS background is programming rather than statistics will find this the least familiar of the five domains. The SAS programming fundamentals exam covers the language rather than the modelling discipline, so a strong programming credential does not close this particular gap.

What does measuring model performance involve?

At 25 percent, about 15 questions, Measure Model Performance is tied with logistic regression as the largest domain on the paper. It asks whether a fitted model performs on data it has not seen, which is a different question from whether it fits the data it was built on.

This is the domain where honest practice and exam practice align most closely. Splitting data for validation, comparing candidate models on a common basis, and reading the curves and statistics that describe discrimination are the things a working analyst does before presenting a model to anyone.

The vocabulary is not SAS-specific, which makes independent material genuinely useful here. Practical model evaluation walkthroughs such as Kaggle’s machine learning course cover the same assessment logic from a different toolchain, and understanding the concept outside SAS makes the SAS output easier to read rather than harder.

For candidates weighing the exam against the effort, the assessment skills are also the most portable thing it certifies. Fitting procedures are tied to SAS; judging whether a model is any good is not, which is part of why analyst roles that name this credential tend to pay around the broader data analyst band reported by sources such as PayScale’s data analyst data rather than a SAS-specific premium.

How should you prepare for A00-240?

Preparation for A00-240 should treat the exam as two halves rather than five domains: the fitting half at 55 percent and the assessment half at 45 percent. SAS’s own recommended route is Statistics 1 for the ANOVA and regression material, followed by Predictive Modeling Using Logistic Regression for the heavier logistic content.

  1. Work out which half you are weaker on, because an analyst who fits models daily and an analyst who evaluates them are missing opposite parts of this paper.
  2. Sweep the ANOVA assumptions material first, covering the central limit theorem, the distribution plots, hypothesis framing, error types and power, since six questions rest on concepts you can settle quickly.
  3. Run PROC GLM and PROC TTEST against the same data and compare what each tells you, then use LSMEANS with PDIFF and both the TUKEY and DUNNETT adjustments so the difference between them is experiential.
  4. Fit multiple linear models with PROC REG and PROC GLM, then examine the residuals with PROC UNIVARIATE rather than accepting the fit statistics at face value.
  5. Give logistic regression the largest single block of study time, and practise converting coefficients into something interpretable rather than stopping at the model output.
  6. Build a candidate input set deliberately, making explicit decisions about missing values, extreme values and categorical representation, then record why you made each one.
  7. Hold data back and measure the model on it, comparing at least two candidate models on the same basis so that assessment becomes a comparison rather than a single verdict.
  8. Finish with timed sets of 60 questions in 110 minutes, using the generous clock to read output carefully rather than to second-guess answers you already know.

Six to ten weeks is realistic for a working SAS analyst, and the range depends almost entirely on how much formal statistics sits behind the day job. Someone who already validates models properly can move quickly; someone who has only ever fitted them should plan for the longer end.

On whether the credential earns its keep, the honest answer depends on the role you are aiming at rather than the exam itself, and this credential and your career is worth thinking through before you book rather than after you pass.

Frequently Asked Questions

How many questions are on the A00-240 exam?

Sixty questions in 110 minutes, which is 110 seconds each. The generous pace reflects how many objectives ask you to interpret output rather than recall a definition.

What is the passing score for the SAS regression and modeling certification?

Sixty eight percent, meaning 41 correct answers out of 60 and a margin of 19. That figure comes from the money site’s syllabus page, since SAS publishes no passing score on its own site.

How much does A00-240 cost?

$180 USD, delivered through Pearson VUE. SAS publishes that price on its own certification page, so it is the one numeric field the vendor and the money site both state.

Which domain carries the most marks?

Two tie at 25 percent each: Logistic Regression and Measure Model Performance. Linear Regression and Prepare Inputs follow at 20 percent each, and ANOVA is smallest at 10 percent.

Is this a statistics exam or a SAS exam?

Both, and the syllabus is specific about the SAS half. It names procedures, statements and options directly, including PROC GLM, PROC TTEST, PROC UNIVARIATE, PROC PLM, LSMEANS with PDIFF and ADJUST, and the HOVTEST option.

Does the exam cover logistic regression more than linear?

Yes, by five points. Logistic carries 25 percent against linear’s 20, which reflects how often business modelling predicts a binary outcome rather than a continuous one.

Is there a prerequisite certification?

None is published. SAS lists no prerequisite for this credential, although the exam assumes you can already write SAS code rather than teaching the language.

What is the difference between Tukey and Dunnett adjustments?

Tukey compares every pair of group means against each other; Dunnett compares each group against a single control. The syllabus names both explicitly under the post hoc objectives.

Does a SAS programming credential prepare you for this one?

Only partly. A programming credential covers the language, while this exam covers the modelling discipline. The input preparation and model measurement domains, worth 45 percent between them, are not language topics at all.

Is A00-240 still current?

Yes. SAS lists the credential in two places on its own certification page, under Most Popular Credentials and under Advanced Analytics, with no retirement notice. SAS also runs a separate Viya credential line, which is a different platform rather than a replacement for this exam.

Conclusion

A00-240 is an even paper wearing an uneven name: 60 questions, 110 minutes, $180, a 68 percent bar, and five domains at 25, 25, 20, 20 and 10. Fitting models accounts for 55 percent of it and everything around the fit accounts for 45.

Read the two assessment domains as one continuous subject, because input preparation is where a model goes wrong and model measurement is where you find out. Give logistic regression the largest block, sweep the ANOVA assumptions quickly rather than deeply, and use the generous clock in the exam to read the output properly instead of rushing a question you could have answered from the table in front of you.

Rating: 0 / 5 (0 votes)

The post SAS Regression and Modeling Certification: Beyond the Fit appeared first on Big Data Prep.

]]>
SnowPro Specialty Native Apps: Inside the NAS-C02 Exam https://www.bigdataprep.com/2026/09/15/snowpro-specialty-native-apps-inside-the-nas-c02-exam/ Tue, 15 Sep 2026 00:00:00 +0000 https://www.bigdataprep.com/?p=16389 Native Apps is the SnowPro Specialty exam that treats you as a software vendor rather than a data professional. Manifests, setup scripts, execution rights, release channels and Marketplace listings carry the paper, and two thirds of the marks sit in the two domains that assume you have actually shipped something.

The post SnowPro Specialty Native Apps: Inside the NAS-C02 Exam appeared first on Big Data Prep.

]]>

Of the three SnowPro Specialty exams, Native Apps is the one that asks you to stop thinking like a data professional and start thinking like a software vendor. Versions, patches, release channels, distribution settings, monetisation, security scanning before a listing goes live: the syllabus reads like a shipping checklist, because that is exactly what it is testing.

NAS-C02 covers the full lifecycle of an application that runs inside somebody else’s Snowflake account. Fifty five questions, 85 minutes, scaled scoring with 750 to pass, and 38 percent of the marks in a single build domain. The data science specialism runs the other way, keeping the work inside the platform, and the Snowflake data scientist certification covers it.

Table of Contents

  1. What does SnowPro Specialty Native Apps actually certify?
  2. The exam code has moved to NAS-C02
  3. How is the NAS-C02 exam delivered and scored?
  4. How are the four domains weighted?
  5. Why does Build carry 38 percent on its own?
  6. What does the exam expect about privileges and execution rights?
  7. Release channels and upgrades are the Manage domain in miniature
  8. Why is Deploy the smallest domain at 11 percent?
  9. Who is NAS-C02 actually written for?
  10. How should you prepare for NAS-C02?
  11. Frequently Asked Questions
  12. Conclusion

What does SnowPro Specialty Native Apps actually certify?

SnowPro Specialty Native Apps certifies that you can design, build, deploy and support applications that run inside a consumer’s Snowflake account using the Snowflake Native App Framework. Snowflake describes it on its own Native Apps certification page as covering native application workloads across their complete lifecycle.

The distinction that matters is ownership of the runtime. A normal Snowflake workload runs in your account, on your data, with your privileges. A native app runs in somebody else’s account, on their data, with privileges they have to agree to grant, and it has to keep working when they upgrade it. Everything unusual about the syllabus follows from that one fact.

It is one of three exams in the SnowPro Specialty series, alongside Snowpark and Gen AI. Those two test capability within Snowflake. This one tests distribution.

The exam code has moved to NAS-C02

Snowflake’s current code for this exam is NAS-C02. The older NAS-C01 path on Snowflake’s own certification site now returns a 404, and the live page is version stamped C02, so the successor is not merely announced but already in place.

This matters because most third party study material has not moved. Courses, practice sets and community write ups are still published under NAS-C01, and some catalogue listings carry the old code alongside a current syllabus. Two practical rules follow:

  • Check any study resource against the four domain names and weightings below rather than against the code on its cover
  • Where a resource predates the C02 page, treat its stated question count, duration and scoring as unverified, because beta era figures circulated widely and differ from the current published ones

The domains themselves did not change name, which is why a lot of C01 material remains substantively useful. It is the specifications, not the subject matter, that need checking.

How is the NAS-C02 exam delivered and scored?

NAS-C02 is 55 questions in 85 minutes, priced at $225 USD, and scored on a scaled range of 0 to 1000 with 750 required to pass. It is delivered through Pearson VUE. Snowflake’s certification page does not render a question count, duration or price to an automated reader, so those figures come from the money site’s published syllabus rather than from Snowflake directly.

Field Value
Credential name SnowPro Specialty: Native Apps
Exam code NAS-C02
Questions 55
Duration 85 minutes
Scoring Scaled 0 to 1000, with 750 to pass
Price $225 USD
Delivery Pearson VUE
Domains Four, all weighted

Scaled scoring is worth understanding rather than converting. A 750 on a 0 to 1000 scale is not the same as answering 75 percent of the items correctly, because items are weighted by difficulty and the scale is calibrated across forms. The practical consequence is that there is no clean number of questions you can afford to lose, which argues for even coverage rather than betting on a strong domain carrying a weak one. Running a NAS-C02 practice test by domain is the reliable way to find out where your coverage is thin.

Eighty five minutes across 55 questions is about 93 seconds each, which is adequate rather than generous given how much of this syllabus is scenario shaped.

How are the four domains weighted?

The NAS-C02 syllabus has four domains, and the weighting is heavily front loaded toward building. Build is 38 percent, Manage is 29 percent, Design is 22 percent and Deploy is 11 percent. Build and Manage together are 67 percent of the exam.

Domain What it covers Weight
Build Snowflake Native Applications Application package components, manifest structure, setup scripts and artifacts; versioned schemas; application logic through stored procedures, external functions, UDFs, UDTFs, Streamlit and Snowpark Container Services; data bundling and in-app sharing; distribution settings; external object references; privilege specification and application roles; event tables and event sharing; test, development and debug modes; version and patch release workflows 38%
Manage Snowflake Native Applications Version release management, release directives and release channels; auto-fulfillment; metadata and usage metrics; consumer upgrade policies and maintenance windows; uninstalled application behaviours; idempotent upgrade code, schema changes and data integrity safeguards 29%
Design Snowflake Native Applications Choosing between Secure Data Sharing, Declarative Native Apps and Snowflake Native Apps; schema separation; framework limitations and constraints; security and privilege strategy including least privilege in manifest requests, provider and consumer boundaries, cross-account security and consumer approval workflows; application role hierarchies; data protection; external access from Snowflake 22%
Deploy Snowflake Native Applications Publishing to Snowflake Marketplace, private against public listings, security scanning and approval; monetisation through subscription, usage and container billing; installation procedures, cross-account installation, multiple instances per consumer, consumer-side event tables; troubleshooting installation failures 11%

Converted to a 55 item paper, that is roughly 21 questions on Build, 16 on Manage, 12 on Design and 6 on Deploy. Deploy is the only domain small enough that a candidate could survive losing most of it, and even then only if everything else is solid.

Why does Build carry 38 percent on its own?

Build is 38 percent because it absorbs almost everything that is specific to the Native App Framework rather than to Snowflake generally. Its objectives run from manifest file structure and setup scripts through versioned schemas, application logic in five different execution styles, data bundling, distribution modes and event configuration, to the release workflow itself.

The four things you build in a Snowflake native app: manifest, setup script, application logic and containers

Read as a list it looks sprawling. Read as a build order it is coherent: you declare the package and its manifest, you write the setup script that constructs the application, you add logic, you decide what data travels with the app and what stays behind a boundary, you wire up events so you can see what is happening in a consumer account you cannot log into, and you version the result.

The five ways application logic can run

The syllabus names stored procedures, external functions, UDFs, UDTFs, Streamlit and Snowpark Container Services as the vehicles for application logic. Questions rarely ask what one of them is; they ask which one fits a stated constraint. Container Services in particular brings its own sub objectives around compute pool settings and service endpoint binding, which is a different mental model from the SQL and Python surface the rest of the domain uses. It assumes you are comfortable with the container image conventions defined by the Open Container Initiative, since that is what a compute pool is ultimately running.

The Native App Framework documentation is the correct depth here, and it is unusually well organised around exactly these components. Where the syllabus mentions Python environment configuration through environment.yml, it is worth knowing how the packaging conventions of the wider Python ecosystem map onto what Snowflake accepts, because the exam assumes that background rather than teaching it.

What does the exam expect about privileges and execution rights?

Privileges appear in both Design and Build, and the execution rights material is the single most distinctive thing on this syllabus. The framework lets application code run either with the rights of the application object’s owner or with a restricted caller’s rights, and the syllabus lists both explicitly along with consumer-side grant workflows and the limitations of the restricted mode.

Owner rights keep provider logic hidden while caller rights run within consumer granted limits

This is not an access control detail. The syllabus says outright that execution rights are how intellectual property is protected in a native app, because they determine what a consumer can see of the provider’s logic. A question describing a provider worried about exposing proprietary code is an execution rights question, however it is dressed.

Alongside it sits a specific list of global privileges the exam names: EXECUTE TASK, EXECUTE MANAGED TASK, CREATE WAREHOUSE, MANAGE WAREHOUSES and CREATE DATABASE, plus configuring IMPORTED PRIVILEGES on databases. These are worth knowing as a list, because they are the ones a manifest actually requests and a consumer actually approves.

Release channels and upgrades are the Manage domain in miniature

Manage is 29 percent and its centre of gravity is release management. The syllabus names release directives, both default and custom, and release channels with the specific examples QA, ALPHA and DEFAULT, along with channel types, monetisation implications, version assignment workflows and the privileges needed to run any of it.

The second half of the domain is the consumer’s side of the same coin: upgrade policies, maintenance windows that block updates during set periods, delayed upgrade schedules with start dates, and control over when an automatic upgrade lands after a directive is issued. A provider pushes; a consumer decides when to catch it.

Underneath both sits the maintenance objective, which asks for idempotent upgrade code, managed database schema changes, data integrity safeguards during upgrades and a versioning strategy. Anyone who has shipped software will recognise this as ordinary release engineering. Anyone whose background is purely analytical will find it the least familiar part of the exam.

Why is Deploy the smallest domain at 11 percent?

Deploy is 11 percent, roughly six questions, and it is small because most of it is procedural rather than architectural. Publishing to the Marketplace, choosing a private or public listing, passing security scanning, configuring billing and installing into a consumer account are steps with correct answers rather than trade-offs with defensible alternatives.

Two parts of it do reward attention despite the low weight. The security scanning objectives name code readability requirements, source map requirements for minified code and dependency vulnerability scanning, which are concrete and easy marks. And monetisation covers subscription and usage based billing, container billing with surcharge configuration, and custom billing events, which is genuinely unfamiliar territory for most engineers.

The rest is installation troubleshooting, which reads naturally once you understand the provider and consumer boundary from the Design domain. If you are working through the Specialty series, our walkthrough of the SnowPro Specialty Gen AI exam covers the sibling credential that shares this exam’s format and scoring model.

Who is NAS-C02 actually written for?

Snowflake states the candidate profile explicitly: one or more years of experience with Snowflake in an enterprise environment, plus six or more months of recent hands-on experience developing native applications. It also names basic knowledge of Snowflake-supported languages such as Python, and basic familiarity with development release cycles such as SDLC.

That second pair is the honest filter. A strong Snowflake practitioner with no software release background will find Manage and half of Build unfamiliar, because those domains are about shipping rather than querying. Conversely a software engineer new to Snowflake will read the release material easily and struggle with schema separation, application roles and the framework’s constraints.

The candidates who pass comfortably have shipped at least one real application package, even a trivial one, into a second account. Six months of that is worth more than any amount of reading, and it is precisely what Snowflake’s own profile asks for. If your Snowflake foundations need work first, the SnowPro Specialty Snowpark route covers the developer surface this exam builds on.

How should you prepare for NAS-C02?

Preparation for NAS-C02 has to be hands-on, because the objectives describe actions rather than facts. Six to eight weeks is realistic for someone who already works in Snowflake but has not shipped an application package. The order below follows the build order rather than the domain order, which is how the material actually makes sense.

  1. Build a trivial application package first, with a manifest and a setup script, and install it into a second account so the provider and consumer boundary stops being abstract
  2. Add application logic in at least three of the named styles, including one Streamlit surface and one Snowpark Container Services component, so compute pools and endpoint binding are concrete
  3. Work through privileges next, requesting them in the manifest, building an application role hierarchy, and running the same procedure under owner rights and then under restricted caller rights to see the difference
  4. Configure an event table and share events back, because observability into an account you cannot enter is a recurring theme across two domains
  5. Version the app, issue a patch, set a release directive and move it across release channels, then upgrade it consumer side under a maintenance policy
  6. Finish with the Deploy material, covering listing settings, security scanning requirements and the billing models, then sit timed sets until 55 scenario questions in 85 minutes is comfortable

Skipping step one is the classic mistake. Almost every scenario in the exam assumes you know what actually happens when an app is installed into an account you do not control, and that is not something reading conveys.

Frequently Asked Questions

What is the current exam code for SnowPro Specialty Native Apps?

NAS-C02. Snowflake’s version stamped certification page carries that code, and the older NAS-C01 path on Snowflake’s own site now returns a 404. Much third party study material still names C01.

How many questions are on the NAS-C02 exam?

Fifty five questions within 85 minutes, according to the money site’s published syllabus. Snowflake’s certification page does not render a question count to an automated reader, so this is not an officially confirmed figure.

What score do you need to pass NAS-C02?

Seven hundred and fifty on a scaled range of 0 to 1000. Because the scale is calibrated rather than a raw percentage, it does not correspond to a fixed number of correct answers.

How much does the SnowPro Specialty Native Apps exam cost?

$225 USD according to the published syllabus. Snowflake’s own FAQ quotes prices for the Core and Advanced series but does not state a Specialty price on the certification page itself.

Which NAS-C02 domain carries the most marks?

Build Snowflake Native Applications at 38 percent, roughly 21 of 55 questions. Manage follows at 29 percent, Design at 22 percent and Deploy at 11 percent.

What experience does Snowflake recommend before taking it?

One or more years with Snowflake in an enterprise environment plus six or more months of recent hands-on native application development. Snowflake also names basic knowledge of a supported language such as Python and of development release cycles such as SDLC.

What is the difference between owner rights and restricted caller’s rights?

They determine whose privileges application code executes with, and therefore how much of the provider’s logic a consumer can observe. The syllabus ties this directly to intellectual property protection, so scenarios about exposing proprietary code are usually execution rights questions.

Does the exam cover Snowpark Container Services?

Yes, inside the Build domain. The objectives name compute pool settings, service endpoint binding requirements and container-based application architecture, and Deploy adds container billing and surcharge configuration.

What are release channels and why do they matter?

Release channels control which consumers receive which version, with QA, ALPHA and DEFAULT named in the syllabus. They sit inside the Manage domain along with release directives, version assignment workflows and the privileges required to operate them.

Is NAS-C02 harder than SnowPro Core?

It is narrower and more specialised rather than simply harder. Core tests breadth across the platform; Native Apps tests one framework in depth and assumes software release experience that Core does not require.

Conclusion

NAS-C02 is the SnowPro Specialty exam for people distributing software rather than analysing data. Fifty five questions, 85 minutes, $225 USD, scaled scoring with 750 to pass, and four domains weighted 38, 29, 22 and 11 percent.

Two things decide the outcome. The first is whether you have actually shipped an application package into an account you do not control, because Build and Manage are 67 percent of the paper and both assume it. The second is whether your study material is current: the exam is now NAS-C02, a great deal of what is published still says NAS-C01, and the specifications are the part that moved. Check the domains, build the app, ship a patch, then book it.

Rating: 0 / 5 (0 votes)

The post SnowPro Specialty Native Apps: Inside the NAS-C02 Exam appeared first on Big Data Prep.

]]>
Google Professional Machine Learning Engineer: The Names Moved https://www.bigdataprep.com/2026/09/10/google-professional-machine-learning-engineer-exam-rewrite/ Thu, 10 Sep 2026 00:00:00 +0000 https://www.bigdataprep.com/?p=16376 The capabilities are the same and the labels are not. Google has confirmed that this exam was updated for the move from Vertex AI to Gemini Enterprise Agent Platform, which means the component names in last year's study notes are the wrong answers now.

The post Google Professional Machine Learning Engineer: The Names Moved appeared first on Big Data Prep.

]]>

If you studied for this exam last year, the product names in your notes have changed. Google says so on the certification page itself: the exam was updated to reflect the transition from Vertex AI to Gemini Enterprise Agent Platform.

That is not a cosmetic edit to a marketing page. It runs through the objectives, so the Workbench, Feature Store, Model Registry, Pipelines and Model Monitoring you are expected to name are all now Agent Platform components. The Google Professional Machine Learning Engineer exam, code GCP-PMLE, is 50 to 60 multiple choice and multiple select questions in two hours at 200 US dollars, across six weighted domains that run from low-code AI solutions through to monitoring what you deployed. This guide sets out where the marks sit, what the rename actually affects, and how much of the paper is operations rather than modelling.

What Does the Google Professional Machine Learning Engineer Exam Cover?

GCP-PMLE covers six domains: architecting low-code AI solutions, collaborating across teams to manage data and models, scaling prototypes into ML models, serving and scaling models, automating and orchestrating ML pipelines, and monitoring AI solutions. It follows a model from an idea in a notebook through to something running in production with alerts on it.

The scope is deliberately wide at both ends. One domain is about building a working model in BigQuery ML or AutoML without writing much code at all, and another is about hyperparameter tuning, distributed training across GPUs and TPUs, and choosing between data and model parallelism. The same paper asks about both.

Generative AI is threaded through, not bolted on

Fine-tuning Gemini models, selecting from Model Garden, evaluating with an LLM as a judge, optimising a Gemini application for cost and latency, and protecting against malicious prompting all appear inside domains that were originally about classical machine learning. This is no longer a predictive-modelling exam with an AI section at the end.

What Are the GCP-PMLE Exam Details?

GCP-PMLE is 50 to 60 multiple choice and multiple select questions in two hours, priced at 200 US dollars and registered through Google CertMetrics. The result is reported as pass or fail, with the money-site syllabus putting the threshold at roughly 70 percent. Google offers the exam in English and Japanese.

Field Value
Exam name Google Professional Machine Learning Engineer
Exam code GCP-PMLE
Questions 50 to 60, multiple choice and multiple select
Duration 120 minutes
Passing score Pass / Fail, approximately 70%
Price $200 USD plus tax where applicable
Languages English and Japanese
Registration Google CertMetrics
Published domains 6, weighted approximately

Two hours across up to 60 questions is two minutes each, which is generous by professional-exam standards. It has to be, because the questions typically describe a business problem, a data shape and a constraint before asking which Google Cloud service you would reach for. Google confirms the length, the fee and the format on its own ML engineer certification page.

Google publishes a renewal process rather than a headline validity figure on that page, so check the renewal guidance directly rather than assuming a term.

What Changed When Vertex AI Became Gemini Enterprise Agent Platform?

Google states plainly that the exam was updated to reflect the transition from Vertex AI to Gemini Enterprise Agent Platform, and points candidates to a new exam guide for the current product list. The capabilities have not moved; the names attached to them have, and the objectives now use the new ones throughout.

Agent Platform components GCP-PMLE expects you to name: Workbench where you build, Registry where versions live, Pipelines where runs repeat, Monitoring where drift shows

Practically, that means the components you have to be able to name are now Agent Platform components: Workbench and Colab Enterprise for notebooks, Feature Store for features, Model Registry for versioning, Pipelines for orchestration, Experiments and ML Metadata for tracking, Model Monitoring for drift, Inference for serving, and Model Garden for choosing a foundation model.

Why this matters more than a rename usually would

Multiple choice and multiple select questions are answered by recognising the right option, and the distractors are other real product names. A candidate who learned the Vertex AI vocabulary is being asked to recognise labels they have never used, under time pressure, and against options that will all look plausible.

The fix is cheap but it is not optional: read the current official exam guide and re-learn the component names before doing anything else. Older courses and books will not have caught up.

How Are the Six GCP-PMLE Domains Weighted?

Scaling prototypes into ML models is the largest at roughly 21 percent, followed by serving and scaling models at about 20 and pipeline automation at about 18. Collaboration and data management takes around 16 percent, and low-code AI solutions and monitoring take about 13 percent each. The weightings are published as approximate rather than exact.

Domain Approximate weight What it asks for
Scaling prototypes into ML models ~21% Choosing model type and product, training and troubleshooting, hyperparameter tuning, fine-tuning foundation models, CPU against GPU against TPU
Serving and scaling models ~20% Batch and online inference, prebuilt and custom containers, Model Registry versioning, A/B and canary rollouts, feature serving, endpoint choice
Automating and orchestrating ML pipelines ~18% End-to-end pipelines, validating data and models, retraining policy, CI/CD/CT deployment
Collaborating to manage data and models ~16% Exploring and preprocessing data, choosing the preprocessing tool by scale, handling personal data, notebooks, experiment tracking
Architecting low-code AI solutions ~13% BigQuery ML and AutoML, industry APIs, model selection from Model Garden, tuning Gemini applications for cost and latency
Monitoring AI solutions ~13% Securing against exfiltration and malicious prompting, responsible AI, explainability, drift and skew monitoring, evaluating generative solutions

Read together, training and serving are about 41 percent and the two smallest domains are the two that bracket the lifecycle. That is a fair description of the job: most of an ML engineer’s difficulty sits between a model that works on a laptop and a model that works for users.

Because the questions are scenario-shaped and the product names have just changed, calibrating against real items is the fastest way to find the gaps. Working GCP-PMLE sample questions surfaces both the reasoning gaps and the vocabulary ones at the same time.

How Much of the Exam Is MLOps Rather Than Modelling?

Roughly a third. Pipeline automation is about 18 percent and monitoring about 13, which is 31 percent before counting the serving domain’s rollout strategies and versioning. An engineer who can build an accurate model but has never automated a retraining run is missing a substantial part of the paper.

The four drift signals on the GCP-PMLE monitoring domain: training serving skew, data drift, concept drift and feature attribution drift

The pipeline objectives are specific about what they mean. Validating data and models. Building and orchestrating pipelines from templates or custom solutions. Keeping preprocessing consistent between training and serving, which is the single most common cause of a model that scores well and behaves badly. Then determining a retraining policy and deploying through continuous integration, delivery and training pipelines.

Monitoring is where security lives

The smallest domain carries the most unexpected content. It covers building secure AI systems against data exfiltration, malicious prompting and oversharing with large language models, and it names Model Armor, safety filters and regular expressions as the controls. It also covers responsible AI and bias monitoring, model explainability, and the classic monitoring quartet of training-serving skew, data drift, concept drift and feature attribution drift.

Anyone whose experience is model building rather than model operations should treat these two domains as the priority, not the afterthought.

Which Open-Source Tools Does the Blueprint Name?

The objectives name PyTorch, sklearn and JAX for model development, XGBoost alongside PyTorch for serving containers, Kubeflow Pipelines and Kubeflow on Google Kubernetes Engine for orchestration, Apache Spark and Dataflow for preprocessing at scale, Ray for distributed work, and Managed Service for Apache Airflow for scheduling.

Tool Where it appears What the question tends to ask
PyTorch, sklearn, JAX Prototyping in notebooks Which framework suits the task and how it runs on the platform
XGBoost Serving from containers Packaging a non-native framework for inference
Kubeflow Pipelines Experiment tracking and orchestration When to use it against the managed Pipelines service
Apache Spark, Dataflow Preprocessing Choosing the preprocessing tool by data scale and complexity
Apache Airflow, Ray Orchestration and distributed compute Scheduling pipelines and scaling training beyond one machine

The recurring question shape is selection rather than syntax. You are given a workload and asked which of BigQuery SQL, Dataflow, Apache Airflow, Spark or an in-memory Python framework fits it, which is a judgement about scale and complexity rather than a coding exercise.

Kubeflow deserves particular attention because it appears twice, once as pipelines for experiment tracking and once running on Kubernetes for training. The Kubeflow project repository is the reference for what it actually does, which matters when the alternative option in a question is Google’s own managed equivalent.

How Should You Prepare for the GCP-PMLE Exam?

Start with the rename, because everything else depends on using the current names. After that, weight the effort toward training, serving and pipelines, which together are close to 60 percent of the paper. Google publishes no mandatory course, and the objectives assume real production experience rather than tutorial completion.

  1. Read the current official exam guide first and relearn the Agent Platform component names, since older courses still teach the Vertex AI vocabulary the questions no longer use.
  2. Work the scaling domain next, the largest at roughly 21 percent, covering model type selection, hyperparameter tuning and when to fine-tune a foundation model instead of training one.
  3. Learn the accelerator decision properly, including when a TPU beats a GPU and how data parallelism differs from model parallelism.
  4. Deploy the same model twice, once for batch inference and once for an online endpoint, then version both in the Model Registry.
  5. Practise a rollout, comparing two model versions with an A/B test and then a canary deployment, because both are named in the serving objectives.
  6. Build an end-to-end pipeline that validates its data, retrains on a policy and deploys through a continuous integration and delivery flow.
  7. Turn on model monitoring and produce all four failure signals deliberately: training-serving skew, data drift, concept drift and feature attribution drift.
  8. Cover the security objectives explicitly, including protection against data exfiltration and malicious prompting, since they sit in the domain most candidates skim.

The GCP-PMLE study companion on this site lists the objectives as a checklist if you want something to tick off as you go.

Who Should Sit This Exam Rather Than a Generative AI Credential?

GCP-PMLE is for people who build and run models. If your work is training, deploying, orchestrating and monitoring, this is the credential that matches it, and the generative AI material is included rather than separate. A leadership-oriented AI credential answers a different question about you entirely.

The dividing line is hands-on responsibility. This exam asks which accelerator to buy time on, how to keep preprocessing consistent between training and serving, and what to do when feature attribution drifts. Those are engineering decisions with operational consequences, and they are hard to answer from a strategy background.

It is also worth noting that generative AI now runs through every domain here, so choosing this exam does not mean choosing classical machine learning over modern work. Our Generative AI Leader guide covers the business-facing alternative if that is closer to your role.

Frequently Asked Questions

How long is the Google Professional Machine Learning Engineer exam?

Two hours, which Google states as the length on its own certification page. The money-site syllabus records the same figure as 120 minutes, and the paper holds 50 to 60 questions.

What is the passing score for GCP-PMLE?

The result is reported as pass or fail. Google publishes no percentage, and the money-site syllabus puts the effective threshold at approximately 70 percent, so plan above that rather than at it.

How much does GCP-PMLE cost?

200 US dollars plus tax where applicable, registered through Google CertMetrics. Foreign exchange and local tax vary the final amount depending on where the exam is booked.

Has the ML engineer exam changed recently?

Yes. Google states that the exam was updated to reflect the transition from Vertex AI to Gemini Enterprise Agent Platform, and directs candidates to a new exam guide for the current product names.

Which GCP-PMLE domain is worth the most?

Scaling prototypes into ML models at roughly 21 percent, closely followed by serving and scaling models at about 20 and pipeline automation at about 18. All the weightings are published as approximate.

Does the exam cover generative AI?

Throughout. Fine-tuning Gemini models, selecting from Model Garden, evaluating with an LLM as a judge and defending against malicious prompting all sit inside domains that also cover classical machine learning.

Is MLOps a big part of GCP-PMLE?

Roughly a third of the paper. Pipeline automation is about 18 percent and monitoring about 13, before counting the rollout strategies and model versioning that sit inside the serving domain.

Which languages is the exam offered in?

English and Japanese. Google lists both on the certification page, and there is no separate regional variant of the exam content.

Do I need to know Kubeflow for the exam?

It is named twice, once as Kubeflow Pipelines for orchestration and experiment tracking and once as Kubeflow on Google Kubernetes Engine for training. Knowing when to use it against the managed alternative is the examinable part.

What monitoring signals does the exam expect?

Training-serving skew, data drift, concept drift and feature attribution drift, along with continuous evaluation metrics for production models and the evaluation of generative solutions.

Conclusion

Before anything else, check whether your study material still uses the right product names. Google has said openly that this exam moved from Vertex AI to Gemini Enterprise Agent Platform, and on a multiple select paper the names are the answers.

Once the vocabulary is current, let the weightings do the planning. Training, serving and pipelines are close to 60 percent between them, so time spent on model selection, accelerator choice, batch against online inference and a real continuous training pipeline pays back faster than anything else. Give the monitoring domain more respect than its 13 percent suggests, because it holds the security material and the four drift signals, and both are easy marks for anyone who has actually turned monitoring on. Then work scenario questions until the product names come back without effort, since two minutes a question is only generous if you are not translating from last year’s names as you read.

Rating: 0 / 5 (0 votes)

The post Google Professional Machine Learning Engineer: The Names Moved appeared first on Big Data Prep.

]]>
The Model Context Protocol Associate Exam Is Half Execution and Security https://www.bigdataprep.com/2026/09/08/model-context-protocol-associate-exam-mcpa/ Tue, 08 Sep 2026 00:00:00 +0000 https://www.bigdataprep.com/?p=16365 On most protocol exams architecture is the biggest domain. On MCPA it is the smallest at 14 percent, while running a tool call and authorising it take half the marks between them. Sixty questions, 120 minutes, 75 percent, and a specification published as dated revisions rather than numbered releases.

The post The Model Context Protocol Associate Exam Is Half Execution and Security appeared first on Big Data Prep.

]]>

The thing this exam grades does not have a version number. It has a date. The Model Context Protocol is published as dated revisions, the current one being 2026-07-28, and that tells you something about the subject before you read a single objective: it is a live specification being worked on in the open, not a product that ships once a year with a release note.

Certifying against a moving standard is an unusual proposition, and the Linux Foundation has answered it by weighting the exam heavily toward the parts that change least. Interactions and Execution takes 26 percent and Security and Governance takes 24, so half the paper sits on how a tool call actually runs and who is allowed to authorise it. This walkthrough covers all five weighted domains, the mechanics, and what the vendor expects you to know before you book.

What Is the Model Context Protocol Associate Exam?

The Model Context Protocol Associate exam, code MCPA, is the Linux Foundation’s associate level credential for the protocol that connects language models to external tools and data. It is a 120 minute online proctored multiple choice exam covering five weighted domains: MCP fundamentals, architecture and components, interactions and execution, security and governance, and use cases and ecosystem.

The protocol itself is the reason the credential exists. Before MCP, every integration between a model and a system was bespoke, which meant an organisation ended up with as many connection styles as it had tools. MCP replaces that with one message format and one set of roles, so a tool built once can be used by any host that speaks the protocol.

The exam is written at the level of someone who has to make that work in practice rather than someone building the protocol itself. It asks what a tool invocation looks like, what happens when one fails, where the trust boundary sits, and who consented to what. Those are integration questions, not research questions.

Which Version of the Spec Does the Exam Follow?

The Model Context Protocol is published as dated revisions rather than numbered releases, and the current revision at the time of writing is 2026-07-28. The Linux Foundation does not pin the MCPA exam to a single dated revision in its published exam information, so the practical guidance is to study the current specification and expect the durable concepts rather than edge case details.

That is less troubling than it sounds, because the five domain names describe structural ideas rather than syntax. Hosts, clients and servers, tool invocation, error handling, trust boundaries and consent are all concepts that survive a revision. What changes between dated revisions tends to be the precise shape of a message or the addition of a capability, and an associate level exam is not the place for that.

The specification is published openly at the dated MCP specification, and the work happens in public on the project repository. Reading a recent set of changes there is the fastest way to see which parts of the protocol are settled and which are still moving.

The certification is valid for two years, which is a sensible match for a standard on this cadence. It is long enough to be worth holding and short enough that the credential still means something about current knowledge when it is renewed.

What Are the MCPA Exam Details?

MCPA carries 60 questions in 120 minutes, costs 250 US dollars and requires 75 percent to pass. It is delivered online and proctored, in multiple choice format, and the purchase includes one retake. There are no prerequisites, and the certification is valid for two years.

Field Value
Exam name Linux Foundation Model Context Protocol Associate
Exam code MCPA
Number of questions 60
Duration 120 minutes
Passing score 75 percent
Price 250 US dollars
Format Online, proctored, multiple choice
Included One retake, certification valid for two years

Two minutes per question is generous by any standard, and it is the most reader friendly mechanic on the paper. This is not a speed test. It is a comprehension test with enough time to read a scenario properly, which fits a subject where the question is usually about what should happen rather than what a command does.

Seventy five percent means 45 correct answers out of 60, a margin of fifteen. Across five domains that is three questions each, which is tighter than it looks once you notice that the two largest domains carry thirty questions between them. Rehearsing against MCPA practice questions is worth doing at full length, because the time allowance changes how the paper feels.

The included retake is genuinely unusual and worth factoring into the price comparison. At 250 US dollars with a second attempt in the box, the effective cost per attempt is lower than several cheaper looking exams that charge again for a resit.

How Is MCPA Weighted Across Five Domains?

MCPA splits into Interactions and Execution at 26 percent, Security and Governance at 24 percent, Use Cases and Ecosystem at 20 percent, MCP Fundamentals at 16 percent and Architecture and Components at 14 percent. The five figures total exactly 100 percent, and the Linux Foundation publishes the same weightings on its own exam page.

Domain Weight What it really covers
Interactions and Execution 26% How a tool call runs, and what happens when it fails
Security and Governance 24% Who is allowed to do what, and how it is recorded
Use Cases and Ecosystem 20% Where MCP is adopted and who owns which part
MCP Fundamentals 16% Purpose, scope and the value of interoperability
Architecture and Components 14% Hosts, clients, servers and structured data

The ordering is the interesting part. On most protocol exams the architecture domain is the biggest, because knowing the components is the foundation of everything else. Here architecture is the smallest at 14 percent, and execution and security take half the paper between them.

That is a deliberate statement about what an associate is expected to be useful for. Knowing that a host talks to a server is assumed background. Knowing what happens when the server returns an error mid invocation, or whether the user actually consented to the action being taken, is the job.

Why Is Interactions and Execution the Largest Domain?

Interactions and Execution is 26 percent of MCPA, roughly sixteen questions, and it is the only domain with four published objectives: interaction patterns and response handling, error handling, the tool invocation lifecycle, and protocol primitives. It is the operational heart of the protocol and the part a working integration lives or dies on.

The four stages of an MCP tool invocation lifecycle from decision to returned result

The tool invocation lifecycle is the concept to get completely solid. A host decides a tool is needed, a client issues the call through the protocol, a server performs the work, and a result travels back into the model’s context. Every question about timing, failure or partial results is a question about which stage of that sequence something went wrong in.

Error handling being its own named objective, rather than a footnote inside response handling, is the clearest signal in the blueprint. Distributed calls fail in ways local function calls do not: the server may be unreachable, the tool may succeed but return something unusable, or the call may time out after the work has already happened. An associate is expected to know what the protocol says should happen in each case.

Protocol primitives are the vocabulary underneath all of it, and they are best learned from the MCP introduction rather than from a summary, because the names are precise and the exam uses them precisely.

How Much of the Exam Is Security and Governance?

Security and Governance is 24 percent of MCPA, roughly fourteen questions, and carries four objectives: trust boundaries, permissions and consent, risk and safety controls, and auditability and observability. Nearly a quarter of an associate level protocol exam is about authorisation and accountability, which is unusual and deliberate.

Trust boundaries are the organising idea. In an MCP deployment the model, the host application, the client and the server are frequently owned by different parties, and each hop is a place where something can be trusted that should not be. Questions here tend to describe an arrangement and ask which component is trusting which, and whether it should.

Permissions and consent are the human half. A model deciding to call a tool is not the same as a user agreeing that it should, and the protocol treats those as separate things. This is the objective most likely to be underestimated by an engineer who thinks of MCP as plumbing, because it is a product design question as much as a technical one.

The risk and safety controls objective connects directly to work being done outside the protocol, and the OWASP LLM application risks is the most useful independent reading for it. Several of the failure modes it describes, particularly excessive agency and insecure plugin design, are exactly what trust boundaries and consent controls exist to prevent.

Auditability and observability close the domain. If a model took an action, the exam expects you to know that the record of it is a design requirement rather than a nice to have.

What Does the Architecture Domain Cover?

Architecture and Components is the smallest domain at 14 percent, roughly eight questions, and covers three things: schemas and structured data, MCP hosts, clients and servers, and the model interaction flow. It is small because it is foundational rather than because it is optional, and the other four domains all assume it.

The three Model Context Protocol roles: host, client and server, and what each one does

The three role names are precise and are frequently confused. A host is the application the user is actually in. A client is the connector inside that host which speaks the protocol. A server is the thing exposing tools or data on the other side. One host can hold several clients, each talking to a different server, and questions that seem to be about topology are usually testing whether those three definitions are clean.

Schemas and structured data are the second half, and they are what make the protocol useful rather than merely possible. A tool describes itself in a structured way so a model can decide whether and how to call it, which is why the vendor lists reading and interpreting server manifests among its recommended skills.

Model interaction flow ties the two together: how a request travels from the user, through the host and client, to the server, and how the result comes back into the model’s context. Get that path straight and the 26 percent execution domain becomes considerably easier.

What Should You Already Know Before Booking?

The Linux Foundation states there are no prerequisites for MCPA, but publishes five recommended competencies: foundational knowledge of JSON-RPC or similar message based protocols, experience interacting with LLM APIs, an understanding of agentic AI concepts, basic literacy in security concepts, and the ability to read and interpret MCP server manifests.

Those five are worth reading as a self assessment rather than as marketing. The exam is not going to test JSON-RPC directly, but it will describe message exchanges in terms that assume you have seen one before. The same is true of the security literacy item, which is not about holding a security credential and is about recognising what a trust boundary means.

The manifest reading skill is the most concrete of the five and the easiest to acquire. Open a few published MCP server manifests, work out what each tool declares about itself, and the architecture and execution domains both become noticeably less abstract. The full list sits on the official MCPA exam page.

For anyone new to the Linux Foundation’s associate tier, the format and pacing are consistent across its credentials, and our walkthrough of the CNPA associate exam shows what that tier expects in a different subject.

How Should You Prepare for the MCPA Exam?

Preparation for MCPA should follow the weighting rather than the blueprint order, because execution and security carry half the paper and architecture carries only 14 percent. Read the current specification once end to end, then build something small, then go back to the two heavy domains with real experience behind you.

  1. Read the current dated revision of the specification once end to end, without stopping to memorise, so the shape of the protocol is familiar before any objective list is opened.
  2. Get the three role names completely clean, host, client and server, and describe an interaction out loud until the path from user to tool and back is automatic.
  3. Open several published MCP server manifests and work out what each declared tool is telling a model about itself, since the vendor lists that skill explicitly.
  4. Build or run a small MCP server, call one of its tools from a host, and then break it on purpose so the error handling objective is something you have watched rather than read.
  5. Work through trust boundaries by drawing the deployment you built and marking every point where one component trusts another, then ask what would happen if that trust were misplaced.
  6. Study permissions and consent as a design question rather than a code question, because the exam separates a model deciding to act from a user agreeing that it should.
  7. Finish on the ecosystem domain, which is 20 percent and largely about roles, adoption and portability, then rehearse the full 60 questions in 120 minutes.

The wider Linux Foundation credential set, including how the associate tier fits together, is collected on our Linux Foundation certification hub.

Frequently Asked Questions

How many questions are on the MCPA exam?

Sixty questions in 120 minutes, which is two minutes per item. It is one of the more generous time allowances at this level.

What is the passing score for the Model Context Protocol Associate exam?

Seventy five percent, which is 45 correct answers out of 60. That leaves a margin of fifteen questions across five domains.

How much does MCPA cost?

250 US dollars, purchased through the Linux Foundation training portal. The price includes one retake.

Are there prerequisites for MCPA?

No. The Linux Foundation states there are no prerequisites, but recommends foundational knowledge of JSON-RPC or similar message based protocols, experience with LLM APIs, an understanding of agentic AI concepts, basic security literacy, and the ability to read MCP server manifests.

Which MCPA domain carries the most weight?

Interactions and Execution at 26 percent, followed by Security and Governance at 24 percent. Together they are half the paper. Use Cases and Ecosystem is 20 percent, MCP Fundamentals 16 percent, and Architecture and Components 14 percent.

How long is the MCPA certification valid?

Two years, which matches the cadence of a specification that is published as dated revisions rather than numbered releases.

What is the difference between an MCP host, client and server?

The host is the application the user is in, the client is the connector inside it that speaks the protocol, and the server is what exposes tools or data on the other side. One host can hold several clients, each connected to a different server.

Is the exam tied to a specific version of the specification?

The Linux Foundation does not pin the exam to a single dated revision in its published exam information. The domain names describe structural concepts that survive revisions, so studying the current specification is the sensible approach.

Does MCPA test security?

Yes, heavily. Security and Governance is 24 percent and covers trust boundaries, permissions and consent, risk and safety controls, and auditability and observability.

Is the exam proctored?

Yes. It is delivered online and proctored, in multiple choice format, through the Linux Foundation training portal.

Conclusion

The Model Context Protocol Associate exam is a protocol credential that spends half its marks on running a call correctly and authorising it properly. Sixty questions in 120 minutes, 250 US dollars with a retake included, 75 percent to pass, valid for two years, and weighted 26, 24, 20, 16 and 14 across five domains.

Study it against the current dated specification rather than a summary, build something small enough to break on purpose, and treat trust boundaries and consent as design questions rather than implementation details. Architecture is the smallest domain, which is the blueprint saying plainly that knowing the components is where this subject starts rather than where it ends.

Rating: 0 / 5 (0 votes)

The post The Model Context Protocol Associate Exam Is Half Execution and Security appeared first on Big Data Prep.

]]>
Retrieval Outweighs Prompting on the Alibaba LLM Exam https://www.bigdataprep.com/2026/09/04/llm-acp-alibaba-rag-heavy-syllabus/ Fri, 04 Sep 2026 00:00:00 +0000 https://www.bigdataprep.com/?p=16342 Sixty correct answers from seventy-five, across six domains weighted within five points of each other. There is no topic you can write off, and the biggest one is retrieval rather than prompting.

The post Retrieval Outweighs Prompting on the Alibaba LLM Exam appeared first on Big Data Prep.

]]>

A user asks a question, the retriever returns three chunks, and every one of them is from the wrong section of the document. Nothing crashed. The model answered confidently. The answer was wrong, and the reason sits somewhere in parsing, chunking, retrieval or reranking. Diagnosing that is the single largest thing the Alibaba Cloud LLM Engineer certification asks you to do.

Retrieval augmented generation is worth 20 percent of LLM-ACP, more than any other domain, and the syllabus names LlamaIndex and the RAGAS evaluation framework outright rather than describing them generically. Add fine-tuning at 16 percent and agent building at another 16 and just over half the exam is about making systems work rather than about explaining what a large language model is. This guide counts the six domains as published, sets out what an 80 percent pass mark actually costs you in wrong answers, and is direct about who the credential travels well for.

What Is the Alibaba Cloud LLM Engineer Certification?

It is the professional-tier credential earned by passing exam LLM-ACP, formally named Alibaba Cloud LLM Engineer (Professional). The exam runs 75 questions in 120 minutes, costs $200 US dollars, and is passed at 80 out of 100 points. Alibaba Cloud aims it at generative AI developers who already have a programming foundation.

That last qualifier is not decoration. The vendor describes the credential as developing the ability to design and implement large language model driven solutions for complex business scenarios, and the objectives back that up: they ask about API parameters, chunking strategies, fine-tuning datasets, agent orchestration and deployment on specific compute services. There is no domain covering what artificial intelligence is.

The exam is offered in English only, and Alibaba Cloud requires fourteen days between any two professional exams, which matters if you are stacking credentials on a deadline. It is also non-refundable, so the $200 is committed at booking rather than at sitting.

Why Is Retrieval the Largest Domain?

Because retrieval is where production LLM applications actually break. At 20 percent, LLM retrieval augmented generation is the biggest single block on the exam, and its objectives run from file parsing and text chunking through retrieval and reranking to two named optimisation strategies and a named evaluation framework.

The retrieval pipeline LLM-ACP tests, running from parse and chunk through retrieve and rerank to evaluation with RAGAS

The technique itself is well documented outside any vendor’s material, and the general retrieval augmented generation overview is worth reading first if the term is new. The syllabus is unusually specific about what it expects you to know.

  • The core components as a pipeline: parsing, chunking, retrieval and reranking, treated as four stages that can each fail independently.
  • Two named retrieval strategies, sentence window retrieval and auto-merging retrieval, which are optimisations rather than defaults.
  • Practical optimisation work: text parsing quality, title rewriting, and table content enhancement, all of which are document-preparation problems rather than model problems.
  • Automated evaluation through the RAGAS framework, so the exam expects you to measure retrieval quality rather than eyeball it.

Two of those names are worth meeting before the exam rather than during it. The pipeline the objectives describe is built on the LlamaIndex project repository, and the evaluation method has its own RAGAS documentation. Neither is Alibaba software, which is a useful signal: the retrieval domain tests transferable engineering rather than a proprietary console.

How Are the Six Domains Weighted?

Retrieval augmented generation leads at 20 percent, followed by LLM application development at 17 percent. Fine-tuning, AI agent applications, and production practices with security compliance sit level at 16 percent each, and prompt engineering closes the syllabus at 15 percent. The six weightings sum to exactly 100.

Domain Weight Approximate questions of 75
LLM retrieval augmented generation 20% 15
LLM application development 17% 13
LLM fine-tuning 16% 12
AI Agent Applications 16% 12
Production Practices and Security Compliance 16% 12
LLM prompt engineering 15% 11

The flatness of that distribution is the point. No domain is dismissible, and the gap between the largest and smallest is five percentage points, or about four questions. An exam weighted this evenly punishes the common strategy of writing off one topic, which is worth knowing before you plan around LLM-ACP exam material rather than after.

What Does an Eighty Percent Pass Mark Demand?

Sixty correct answers out of 75, and no more than fifteen wrong. That is a materially harder bar than most vendor credentials set, and against a flat six-domain distribution it means you cannot afford to lose a whole domain: the smallest domain is worth roughly eleven questions, which would consume nearly the entire error budget on its own.

The LLM-ACP eighty percent bar shown as sixty correct answers needed, fifteen wrong allowed and ninety six seconds per question

The arithmetic is worth sitting with. Fifteen permitted mistakes across six domains averages two and a half per domain. Drop the retrieval domain entirely and you are fifteen down before the other five have been marked. There is no combination of strengths that rescues one abandoned topic.

Pace is the second constraint. One hundred and twenty minutes across 75 questions is 96 seconds each, which is comfortable for recall and tight for a scenario that names three deployment services and asks which balances cost against latency. Build the habit of committing to an answer and moving on, because the marginal value of a second reading is low when the bar is 80 percent.

Which Named Tools Does the Syllabus Assume?

The objectives name specific technologies rather than describing capabilities in the abstract, and there are more of them than a professional syllabus usually carries. Three come from the open ecosystem and four are Alibaba Cloud services, which tells you how the exam splits between transferable and vendor-specific knowledge.

Named in the syllabus Domain it appears in What you are expected to do with it
LlamaIndex Retrieval augmented generation Build a RAG pipeline and understand its parsing, chunking, retrieval and reranking stages
RAGAS Retrieval augmented generation Evaluate a retrieval system automatically rather than by inspection
vLLM Production practices Deploy a fine-tuned model for serving
Model Studio Agents, and production practices Build agents through the Model API and deploy fine-tuned models
Elastic Compute Service Production practices Host a fine-tuned model on general compute
Platform for AI Production practices Deploy through the managed machine learning platform
Function Compute Production practices Publish an AI assistant serverlessly

Read that table as a study checklist. Three of the seven, LlamaIndex, RAGAS and vLLM, are open projects you can install and use today at no cost, which makes the largest domain the cheapest one to rehearse properly.

What Are the LLM-ACP Exam Details?

LLM-ACP is 75 questions in 120 minutes for $200 US dollars, passed at 80 out of 100 points, scheduled through Pearson VUE and delivered either online or at a test centre. It is available in English only and cannot be refunded once booked.

Detail Value
Exam name Alibaba Cloud LLM Engineer (Professional)
Exam code LLM-ACP
Questions 75
Duration 120 minutes
Passing score 80 of 100
Price $200 USD, non-refundable
Language English only
Delivery Online or at a test centre, chosen at booking
Retake spacing 14 days between any two professional exams
Scheduling Pearson VUE

The fourteen-day rule deserves a note. It applies between professional exams generally, not only to retakes of the same paper, so a failed attempt costs two weeks before a second sitting and blocks any other professional credential in that window. Alibaba Cloud sets all of this out on the official LLM Engineer certification page, which also confirms the recommended course runs to three chapters and fifteen lessons.

Is This Credential Worth It Outside Its Home Markets?

Partly, and the split is unusually clean. Roughly half the syllabus covers work that transfers anywhere: retrieval pipelines, chunking strategy, fine-tuning method, prompt frameworks and agent orchestration are the same problems on any cloud. The other half names Alibaba Cloud services you will only use if your organisation runs there.

For engineers working in markets where Alibaba Cloud has a real footprint, the calculation is simple and favourable. For everyone else, the honest framing is that the retrieval, fine-tuning and prompt domains are worth studying regardless, while the production domain is a vendor tour that will not appear on your next project.

There is also a recognition question. The credential has very little independent editorial coverage, which is precisely why an article that counts the weightings is useful, but it also means a hiring manager outside the vendor’s ecosystem may not recognise the name. That is a real cost against $200, and it should be weighed rather than dismissed.

Readers coming from a data-platform background rather than an application one should look at the adjacent associate credential first: our guide to whether the ACA Data Engineer credential fits covers who that tier actually suits.

How Should You Prepare for LLM-ACP?

Build the retrieval pipeline first and study everything else around it. The largest domain is also the one you can rehearse for free with open tooling, and doing so teaches the chunking, parsing and evaluation vocabulary that the fine-tuning and agent domains then build on.

  1. Build a working retrieval pipeline with LlamaIndex over your own documents, so that parsing, chunking, retrieval and reranking become four stages you have debugged rather than four words you have read.
  2. Measure that pipeline with RAGAS rather than by reading its answers, because the syllabus asks for automated evaluation and the habit of inspecting outputs by hand will not survive the questions.
  3. Work the two named retrieval optimisations, sentence window and auto-merging, against the same corpus so the difference between them is something you have observed rather than memorised.
  4. Move to fine-tuning next, concentrating on dataset construction and parameter settings, which is where the objectives put their weight rather than on algorithm theory.
  5. Build one agent through Model Studio and extend it into a multi-step workflow, since agents and production practices together carry 32 percent and share the same deployment vocabulary.
  6. Finish with prompt engineering and application development as a confidence pass, then rehearse against the clock at 96 seconds per question with a target of 60 correct.

The order is deliberate rather than conventional. Most preparation plans open with prompt engineering because it feels like the entry point; it is in fact the smallest domain at 15 percent, and starting there spends your freshest study time on your cheapest marks.

Where It Sits Against the Associate Credentials

LLM-ACP is the professional tier of Alibaba Cloud’s language model track, sitting above an associate LLM credential and alongside the wider associate programme covering cloud engineering, data engineering and generative AI. The professional tier is where the syllabus stops describing and starts building.

The gap between the tiers is visible in the objective verbs. Associate material asks you to identify and describe; this exam asks you to construct a RAG pipeline, evaluate it, fine-tune a model and deploy it across three named compute services. Anyone who has not yet built one of these systems end to end will find the associate route a better first step, and our generative AI engineer study guide covers what that tier expects.

For candidates already deploying LLM applications, the professional exam is the one that matches the work. The fourteen-day spacing rule between professional exams is worth planning around if you intend to take more than one in the same quarter.

Frequently Asked Questions

What is the Alibaba Cloud LLM Engineer certification?

A professional-tier credential earned by passing exam LLM-ACP. It validates the ability to build, fine-tune, deploy and evaluate large language model applications on Alibaba Cloud, and targets developers who already have a programming foundation.

How many questions are on the LLM-ACP exam?

Seventy-five, with a 120-minute limit. That works out at roughly 96 seconds per question, which is comfortable for recall items and tight for scenarios naming several deployment services.

What is the passing score for LLM-ACP?

Eighty out of 100 points, which means 60 correct answers from 75 and a budget of fifteen wrong. That is a higher bar than most vendor credentials set at professional level.

How much does the LLM-ACP exam cost?

Two hundred US dollars, and the fee is non-refundable once booked. Scheduling runs through Pearson VUE, with the choice of an online or a test-centre sitting made at booking time.

Which LLM-ACP domain is the largest?

Retrieval augmented generation at 20 percent, roughly 15 of the 75 questions. It covers parsing, chunking, retrieval and reranking, two named retrieval optimisations, and automated evaluation with RAGAS.

Does the exam cover LlamaIndex?

Yes, by name. The retrieval domain asks candidates to build RAG with LlamaIndex and to understand its core components, which makes it one of the few open-source tools examined directly by a vendor credential.

What language is the LLM-ACP exam available in?

English only, according to Alibaba Cloud’s own exam overview. That is narrower than most cloud credentials, which typically offer several languages at professional level.

How long must you wait between Alibaba Cloud professional exams?

Fourteen days. The rule applies between any two professional exams rather than only to retakes, so a failed attempt also blocks a different professional credential during that window.

Is prompt engineering a big part of LLM-ACP?

It is the smallest domain at 15 percent, roughly 11 of the 75 questions. It covers prompt frameworks, separators and templates, the role of the system prompt, and applied tasks such as batch intent classification.

Do you need an associate credential before LLM-ACP?

No formal prerequisite exists. Alibaba Cloud recommends a programming foundation rather than a prior credential, though candidates who have not built an LLM application end to end will find the associate tier a more realistic starting point.

Conclusion

The Alibaba Cloud LLM Engineer certification is more technical than its low profile suggests. Retrieval augmented generation at 20 percent, fine-tuning and agents at 16 percent each, named open tooling in the objectives, and an 80 percent pass mark that leaves room for only fifteen wrong answers across six evenly weighted domains.

Start with the retrieval pipeline, because it is the largest domain and the one you can build for free before spending anything on the exam. If that work feels like your day job, the credential describes what you already do. If it feels like a new subject, the associate tier is the better first purchase.

Rating: 5 / 5 (1 votes)

The post Retrieval Outweighs Prompting on the Alibaba LLM Exam appeared first on Big Data Prep.

]]>
Nothing Is Left Vague on the SAS Programming Fundamentals Exam https://www.bigdataprep.com/2026/09/03/sas-a00-215-programming-fundamentals-named-syntax/ Thu, 03 Sep 2026 00:00:00 +0000 https://www.bigdataprep.com/?p=16329 Most blueprints leave depth to the reader. This one lists CATX, SCAN, MDY and OBS= by name, which turns preparation into a checklist you can genuinely finish.

The post Nothing Is Left Vague on the SAS Programming Fundamentals Exam appeared first on Big Data Prep.

]]>

Sixteen functions. Four data set options. Five procedures. That is not a summary of the SAS programming fundamentals exam, it is what the published syllabus actually lists, by name, in the objectives themselves. A00-215 tells you to use UPCASE, PROPCASE, SUBSTR, SCAN, FIND, LENGTH and CATX. It tells you to use MONTH, DAY, YEAR, TODAY and MDY. It tells you ROUND, INT, MEAN and SUM. It names DROP=, KEEP=, RENAME= and OBS=.

Very few certification blueprints are that literal, and it changes what preparation looks like. Instead of guessing what depth is expected, you can build a checklist and work down it until nothing is unfamiliar. The exam gives you 120 minutes for 60 to 65 questions, asks for 68 percent, and costs $120, and it publishes no weightings at all across its seven topics. This guide goes through the whole named list, explains where the marks are likely to concentrate anyway, and shows what the absence of weightings means for a study plan.

Why Does the A00-215 Syllabus Name Individual Functions?

Because SAS is examining a fixed language surface rather than a shifting product. The A00-215 objectives repeatedly say “use” followed by an explicit list, so the exam commits to a bounded set of syntax rather than an open-ended one. That makes the syllabus unusually testable against, and it means a candidate can genuinely know when preparation is finished.

Compare that with a cloud certification, where an objective might read “configure identity and access” and leave the depth entirely to the reader. Here the objective reads “Use Character Functions: UPCASE, PROPCASE, SUBSTR, SCAN, FIND, LENGTH, CATX”. There is no ambiguity about scope, only about how deep each item goes.

The practical consequence is that a checklist beats a course for the final stretch. Read the objectives, list every named item, and mark each one as known, shaky or unseen. Anything still marked unseen a week before the exam is a genuine risk, and there are not many of them.

The Seven Topics, and Why None of Them Carries a Weight

A00-215 publishes seven topics and no percentage against any of them. They run from fundamental SAS concepts and log reading, through exploring data sets, two separate DATA step topics, report generation with PROC steps, utility procedures, and finally importing and exporting non-SAS files.

Topic What it asks you to demonstrate
Fundamental SAS Concepts Rules for DATA and PROC steps, rules for SAS statements including global statements, interpreting the log, and telling syntax errors from logic errors including using PUTLOG
Explore SAS Data Sets Naming conventions, character and numeric variable types, creating and manipulating date values, missing data, the LIBNAME statement, PROC CONTENTS, and the DROP=, KEEP=, RENAME= and OBS= data set options
Using the DATA Step to Access SAS Data Sets DATA and SET statements, MERGE and BY for horizontal combining, the IN= option, SET for vertical combining, compilation and execution including the program data vector, and subsetting with WHERE, IF, DROP and KEEP
Using the DATA Step to Manipulate Data Creating and updating variables, character, date, truncation and descriptive statistics functions, conditional processing, controlling output, accumulating variables with SUM and BY group FIRST. and LAST., iterative DO loops, and permanent attributes via FORMAT and LABEL
Generate Reports Using PROC Steps PROC PRINT, PROC MEANS and PROC FREQ with their named options, plus TITLE, FOOTNOTE, temporary FORMAT and LABEL, and WHERE for subsetting
Use Utility Procedures PROC SORT with OUT=, BY and DESCENDING, and PROC FORMAT with the VALUE statement and the OTHER keyword
Import and Export non-SAS files PROC IMPORT and PROC EXPORT for CSV, the LIBNAME statement with the XLSX engine, and ODS to PDF, RTF and EXCEL with the FILE= and STYLE= options

The absence of weightings is not an oversight. It means no topic can be discounted on arithmetic, so coverage has to be even. It also means the only honest signal about emphasis is the length of each objective, and by that measure the two DATA step topics dwarf everything else. Once the language surface is secure, the statistical track follows, and the SAS regression modeling certification is where that begins.

What Is the SAS Programming Fundamentals Exam Format?

A00-215 presents 60 to 65 questions with a 120-minute limit and a 68 percent passing score, for $120 US dollars. The exam is administered by SAS and Pearson VUE. Two hours across roughly 62 questions is close to two minutes each, which is generous for a paper at this level.

Detail Value
Exam name SAS 9.4 Programming Fundamentals
Credential SAS Certified Associate: Programming Fundamentals Using SAS 9.4
Exam code A00-215
Questions 60 to 65
Duration 120 minutes
Passing score 68%
Price $120 USD
Registration Pearson VUE
Recommended training SAS Programming 1: Essentials and SAS Programming 2: Data Manipulation Techniques

One detail from the SAS credential page deserves attention: the paper is not purely multiple choice. SAS describes it as multiple choice and short-answer questions, which means some items expect you to supply a value rather than recognise one. That changes revision, because recognising correct syntax and producing it are different skills.

Sixty-eight percent of roughly 62 questions is about 42 correct answers, leaving a margin of 20. That is workable, but it is tighter than the 60 percent bars common at associate level, so the even-coverage requirement bites harder than it first appears.

Which Functions Does the Exam Name by Hand?

Sixteen, split into four groups by purpose. The syllabus lists character functions, date functions, truncation functions and descriptive statistics functions as four separate bullet points, each with its own explicit membership. Learn them by group and the structure of the questions becomes predictable.

Four cards grouping the sixteen functions named in the A00-215 syllabus by purpose
  • Character functions: UPCASE, PROPCASE, SUBSTR, SCAN, FIND, LENGTH, CATX
  • Date functions: MONTH, DAY, YEAR, TODAY, MDY
  • Truncation functions: ROUND, INT
  • Descriptive statistics functions: MEAN, SUM

The character group is the one that rewards real practice. SCAN and FIND both search, but they answer different questions, and CATX concatenates while stripping and inserting separators in a way that trips people who reach for it once a year. SUBSTR appears in almost every SAS exam at every level, and it is worth being able to write it without hesitation.

The date group is smaller but connects to a separate objective about how SAS stores date values and how date formats control display. Those two ideas, storage and display, sit behind more questions than the function list alone suggests. Getting them straight early makes several other objectives easier.

Working through A00-215 sample questions function by function is the most efficient way to find which of the sixteen you can recognise but not reproduce, which matters given the short-answer format.

What Does the DATA Step Section Expect You to Do?

More than any other part of the syllabus. Two of the seven topics are DATA step topics, and between them they carry the longest objective lists on the paper. The first covers reading and combining data sets, the second covers manipulating what you have read.

Four cards showing the WHERE, IF, DROP and KEEP routes to subsetting a SAS data set and when each acts

Combining and subsetting

MERGE and BY combine horizontally, SET combines vertically, and the IN= option on MERGE controls which records survive. Subsetting appears four ways, and the exam expects you to know which happens when: WHERE subsets on input, IF subsets during processing, DROP and KEEP statements subset at output, and the DROP= and KEEP= options subset at both input and output.

The program data vector

The objective asks you to describe how the program data vector is created, how the LENGTH statement changes its default behaviour, and how a DATA step iterates. This is the most conceptual material on the exam and the part that separates people who write SAS from people who understand it.

Accumulating and controlling output

The SUM statement builds a running total, and BY group processing with FIRST. and LAST. resets it per group. Alongside that sit the OUTPUT statement for directing and timing output, iterative DO loops, and permanent attributes assigned with FORMAT and LABEL. None of it is difficult in isolation; the questions come from combining two of them.

Which PROC Steps Are Examinable?

Five, and the syllabus names the specific options for each. PROC PRINT, PROC MEANS and PROC FREQ handle reporting. PROC SORT and PROC FORMAT are the utility procedures. Nothing else appears in the objectives, which makes this the easiest part of the syllabus to bound.

Procedure Named options and statements
PROC PRINT LABEL and NOOBS options, VAR statement
PROC MEANS MAXDEC= option, VAR and CLASS statements
PROC FREQ ORDER= option, TABLES for one-way and two-way, NOCUM and NOPERCENT, CROSSLIST
PROC SORT OUT= option, BY statement, DESCENDING option
PROC FORMAT VALUE statement, OTHER keyword for missing values

PROC FREQ carries the most named options, which is a reasonable proxy for how much attention it gets. Two-way tables with CROSSLIST are the least familiar item for most candidates, and worth an hour on their own. Report enhancement through TITLE, FOOTNOTE and temporary FORMAT and LABEL statements sits alongside these and is quick to learn.

How Much Import and Export Work Is on the Paper?

A full topic of it. The seventh topic covers moving data between SAS and the outside world: PROC IMPORT and PROC EXPORT for CSV files, the LIBNAME statement with the XLSX engine for Excel, and the Output Delivery System for sending reports to PDF, RTF and EXCEL destinations using the FILE= and STYLE= options.

This is the topic most likely to be under-studied, because it feels peripheral next to the DATA step. With no published weightings, though, there is no evidence it carries fewer questions than anything else, and it contains the fewest objectives, which makes it the cheapest topic per hour invested.

The ODS objectives are worth particular care. Knowing that a destination is opened and closed around the procedure that produces the output, and what STYLE= actually changes, covers the realistic question shapes without needing to memorise every available style.

Does the SAS Version Actually Matter?

Yes, and SAS is specific about it. The exam is based on SAS 9.4 M5, and the credential name carries the version too: SAS Certified Associate: Programming Fundamentals Using SAS 9.4. Studying against a much newer Viya-oriented resource risks meeting syntax and interfaces the exam does not assess.

That version pinning is also why the named function list stays stable. The Base SAS language surface described in these objectives has been consistent for a long time, which is unusual in certification and is part of why a checklist approach works here at all. Background on how the language is structured is set out in the SAS language overview if the DATA step and PROC step split is new to you.

For anyone weighing the credential against the effort, an associate programming certification remains a recognisable marker in analytics hiring, and published ranges for SAS programmer pay give a realistic picture of where the skill sits in the market. The credential is a starting signal rather than a senior one, which matches its price and its bar.

Version pinning also decides what comes next. The tier directly above is the performance-based Base Programming exam, and the A00-231 Base Programming route reuses almost everything learned here while moving the assessment into a live environment.

How Do You Study a Syllabus With No Weightings?

By using objective length as the only available proxy for emphasis, and by covering everything rather than optimising. With seven unweighted topics and a 68 percent bar, the safe assumption is that any topic can supply enough questions to matter, so the plan below works through them in dependency order rather than syllabus order.

  1. Begin with fundamental SAS concepts and log reading, because every later topic assumes you can tell a syntax error from a logic error and know what PUTLOG is for.
  2. Cover data set exploration next, including variable types, how SAS stores date values, missing data, LIBNAME, PROC CONTENTS and the four data set options.
  3. Give the two DATA step topics the largest share of your time, working MERGE, SET, the IN= option, the four subsetting routes and the program data vector until each is automatic.
  4. Build the named function checklist by group, then practise writing each of the sixteen rather than recognising them, since some items are short answer.
  5. Work the five procedures with their named options, spending extra time on two-way PROC FREQ tables with CROSSLIST.
  6. Finish with import, export and ODS, which is the shortest topic and the cheapest marks per hour on the paper.

Six to eight weeks suits someone new to SAS working alongside a job, and three or four for a candidate who already writes DATA steps but has never sat a SAS exam. Build the checklist in week one so the remaining weeks have a target to shrink.

One practical note on cost: SAS offers free or discounted training through its academic programmes, which is worth checking before paying for courses. The earlier programming fundamentals overview covers why candidates pursue this credential in the first place.

Frequently Asked Questions

How many questions are on the A00-215 exam?

Sixty to sixty-five questions with a 120-minute limit, which is close to two minutes each. SAS describes them as multiple choice and short-answer items rather than multiple choice alone.

What is the passing score for SAS Programming Fundamentals?

Sixty-eight percent. On roughly 62 questions that means about 42 correct answers, leaving a margin of around 20, which is tighter than the 60 percent bars common at associate level.

How much does the A00-215 exam cost?

One hundred and twenty US dollars worldwide. That makes it one of the cheaper vendor certifications available, and the exam is administered by SAS together with Pearson VUE.

Which SAS version is the exam based on?

SAS 9.4 M5. The credential name carries the version as well, so studying against much newer Viya material risks covering interfaces the exam does not assess.

Do the A00-215 topics have percentage weightings?

No. Seven topics are published with no weights against any of them, which means coverage has to be even and no topic can safely be discounted on arithmetic.

Which functions do you need to know?

Sixteen named functions in four groups: UPCASE, PROPCASE, SUBSTR, SCAN, FIND, LENGTH and CATX; MONTH, DAY, YEAR, TODAY and MDY; ROUND and INT; MEAN and SUM.

Does the exam require writing code by hand?

Partly. Because SAS includes short-answer items alongside multiple choice, some questions expect you to supply a value or a piece of syntax rather than pick one from a list.

What is the difference between A00-215 and A00-231?

A00-215 is the Programming Fundamentals associate exam. A00-231 is the Base Programming performance-based exam, which sits above it and expects you to work in a live environment.

Which training does SAS recommend?

SAS Programming 1: Essentials and SAS Programming 2: Data Manipulation Techniques. SAS also offers free or discounted training routes through its academic programmes, which is worth checking first.

How long does preparation usually take?

Six to eight weeks for someone new to SAS studying alongside a job, and three or four weeks for a candidate who already writes DATA steps but has never sat a SAS certification.

Conclusion

A00-215 is one of the few certification syllabuses you can genuinely finish. Sixteen named functions, four data set options, five procedures with their specific options, four ways to subset, and a program data vector to understand. Nothing is left to inference, and nothing is hidden behind a vague objective.

What the syllabus does not give you is a weighting map, so the plan has to be even rather than clever: fundamentals first, then the two DATA step topics for the bulk of the time, then the function checklist, the procedures and finally the short import and export topic. Once every named item on that list is something you can write rather than merely recognise, sample items across all seven topics will confirm whether 42 correct answers is within reach.

Rating: 0 / 5 (0 votes)

The post Nothing Is Left Vague on the SAS Programming Fundamentals Exam appeared first on Big Data Prep.

]]>
Generative AI Leader Certification: Google’s AI Exam That Never Opens a Console https://www.bigdataprep.com/2026/08/29/generative-ai-leader-certification-gcp-gail/ Sat, 29 Aug 2026 00:00:00 +0000 https://www.bigdataprep.com/?p=16293 Every objective on GCP-GAIL starts with describe or identify, never configure. A working read of the four domains, the twenty five named Google products, and who the credential is really for.

The post Generative AI Leader Certification: Google’s AI Exam That Never Opens a Console appeared first on Big Data Prep.

]]>

Almost every Google Cloud certification assumes you will open a console. The Generative AI Leader certification, exam code GCP-GAIL, assumes you will not. It is 90 minutes, 50 to 60 multiple choice questions, $99, and a pass or fail result at roughly 70 percent, and there is not a single command, query or configuration task anywhere in the syllabus.

What there is instead is a vocabulary test with business consequences. Four domains, weighted 30, 35, 20 and 15 percent, covering what generative AI actually is, what Google sells in that space, how to make a model’s output better, and how to decide whether a gen AI project is worth doing at all. The largest domain is essentially a product catalogue, and it names around twenty five Google services by name. This article groups them into something learnable, walks all four domains, and sets out how to prepare for an exam you cannot practise by building anything. Readers who want the opposite, an exam about running models on real infrastructure, can compare it with the Nutanix NCP-AI certification.

What Does the Generative AI Leader Certification Actually Certify?

It certifies business-level fluency in generative AI, not the ability to build with it. GCP-GAIL asks whether you can define the core concepts, tell Google’s gen AI products apart and say what each is for, describe the techniques that improve model output, and reason about whether and how an organisation should adopt gen AI at all. The building side is examined instead by the Google machine learning engineer, which expects console work throughout.

Google was explicit about this when the credential launched. It arrived in May 2025 as a first of its kind credential, and the intended audience was stated plainly.

“It’s specifically designed for non-technical professionals like managers, strategists and leaders to give them a foundational understanding of AI and how it’s used.”

Erin Rifkin, Managing Director, Google Cloud Learning Services

What that means for how you read the syllabus

Every objective in this exam starts with a verb like describe, identify, define, recognise or explain. None of them says configure, deploy, implement or troubleshoot. If you come from an engineering background that sounds easy, and it usually is, but it also means precision matters more than usual: the difference between grounding and fine-tuning is a definitional distinction here, not a design decision you can reason your way to under pressure. Where this credential stops at describing, the LLM-ACP syllabus expects you to build and tune the retrieval pipeline yourself.

How Are the Four GCP-GAIL Domains Weighted?

Google publishes approximate weightings for all four, which makes planning straightforward. Google Cloud’s own gen AI offerings take the largest share at around 35 percent, the fundamentals take around 30, techniques to improve model output take around 20, and business strategy takes around 15. On a 50 to 60 question paper that is roughly 18 to 21 questions about Google products alone.

Domain Weight Approximate questions What it is really about
Google Cloud’s gen AI offerings ~35% 18 to 21 Naming Google’s products and saying what each one is for
Fundamentals of gen AI ~30% 15 to 18 Core concepts, the ML lifecycle, data types, choosing a foundation model
Techniques to improve gen AI model output ~20% 10 to 12 Model limitations, prompting, grounding, RAG, sampling parameters
Business strategies for a successful gen AI solution ~15% 7 to 9 Adoption steps, measuring impact, secure AI, responsible AI

Read that top row carefully, because it is the thing candidates underestimate. Roughly a third of this exam is Google’s own catalogue, and no amount of general AI literacy will substitute for knowing which product is the agent builder and which is the search one. The Google Generative AI Leader resources collected for this exam are the fastest way to find out whether you can already tell them apart.

What Sits Inside the Gen AI Fundamentals Domain?

Around 30 percent of the paper, covering definitions, learning approaches, the machine learning lifecycle, how to choose a foundation model, the data types that feed gen AI, the layers of the gen AI landscape, and Google’s own model families. It is the domain that rewards careful reading of the syllabus, because the terms it asks you to define are listed explicitly.

Three ways the GCP-GAIL syllabus improves generative AI output: grounding, retrieval augmented generation and tuning

The definitional list is long and specific: artificial intelligence, natural language processing, machine learning, generative AI, foundation models, multimodal foundation models, diffusion models, prompt tuning, prompt engineering and large language models. Note that prompt tuning and prompt engineering are named separately. They are not synonyms, and a question can turn on that alone.

The lifecycle, and a Google tool for every stage

The syllabus asks for the stages of the machine learning lifecycle, data ingestion, data preparation, model training, model deployment and model management, and then asks which Google Cloud tool serves each one. That pairing is the giveaway: this domain and the offerings domain are designed to interlock, so learning the lifecycle without the tool names leaves half the marks on the table.

Model selection is framed as a business decision rather than a technical one. You choose a foundation model on modality, context window, security, availability and reliability, cost, performance and how far it can be fine-tuned or customised. Four Google model families are named directly and you should be able to say what each is for: Gemini, Gemma, Imagen and Veo.

Data literacy closes the domain. Structured against unstructured, labelled against unlabelled, and the six named characteristics of data quality and accessibility: completeness, consistency, relevance, availability, cost and format. The five layers of the gen AI landscape, infrastructure, models, platforms, agents and applications, are the mental map the rest of the exam hangs on.

Why Do Google’s Own Offerings Take the Largest Share?

Because this is a Google credential aimed at people who will recommend Google products, and the syllabus is unapologetic about it. At around 35 percent it is the biggest domain, and it names roughly twenty five distinct services and APIs. The only way through is to group them by what they do rather than to memorise a flat list.

Four useful groups. First, the platform: Vertex AI Platform with Model Garden and AutoML, Vertex AI Search, Vertex AI Agent Builder, and the choice between Vertex AI Studio and Google AI Studio. Second, the ready-made assistants: the Gemini app, Gemini Advanced with Gems, Gemini for Google Workspace, and Google Agentspace with the Cloud NotebookLM API and multimodal search. Third, the customer-facing set: the Customer Engagement Suite, meaning Conversational Agents, Agent Assist, Conversational Insights and Contact Center as a Service. Fourth, the API toolbox agents call: Speech-to-Text, Text-to-Speech, Translation, Document Translation, Document AI, Cloud Vision, Cloud Video Intelligence and Natural Language, alongside Cloud Storage, databases, Cloud Functions and Cloud Run.

The infrastructure and positioning objectives

Underneath the products sit the arguments. Google’s AI-first approach, the enterprise-ready platform described as responsible, secure, private, reliable and scalable, the open approach, and the AI-optimised infrastructure: the hypercomputer, custom-designed TPUs, GPUs and data centres. Two further objectives cover data control and democratisation through low-code and no-code tools, pre-trained models and APIs.

These read like marketing, and in a sense they are, but they are examinable marketing. The practical advice is to learn one concrete example for each claim rather than the claim itself, because the questions tend to be scenario-shaped: a business need is described and you pick the offering that fits.

What Techniques Does the Exam Expect You to Name?

Around 20 percent of the paper, and it is the most technical-feeling domain in an otherwise non-technical exam. It covers foundation model limitations, the recommended practices that address them, continuous monitoring, prompt engineering, grounding, retrieval-augmented generation, and the sampling parameters that control model behaviour.

Four prompting techniques named in the GCP-GAIL syllabus: zero shot, few shot, role prompting and prompt chaining

Start with the limitations, because everything else is a response to one of them: data dependency, the knowledge cutoff, bias, fairness, hallucinations and edge cases. The remedies named are grounding, retrieval-augmented generation, prompt engineering, fine-tuning and human in the loop. Being able to match a remedy to a limitation is the single most testable skill in the domain.

Prompting and grounding, named technique by technique

The prompting list is explicit: zero-shot, one-shot, few-shot, role prompting and prompt chaining, then chain-of-thought and ReAct as the advanced pair. Grounding is split three ways by data source, first-party enterprise data, third-party data and world data, and Google’s grounding offerings are named as prebuilt RAG with Vertex AI Search, RAG APIs and grounding with Google Search.

Sampling parameters close the domain and are pure recall: token count, temperature, top-p or nucleus sampling, safety settings and output length. Know what raising temperature does before you sit down. Monitoring is the quieter half of this domain and is worth the same marks: automatic model upgrades, key performance indicators, security patches, versioning, performance tracking, drift monitoring and the Vertex AI Feature Store.

Business Strategy, Secure AI and Responsible AI

The smallest domain at around 15 percent, and the one that sounds softest while asking some of the most concrete questions. It covers choosing a gen AI solution for a business need, integrating it into an organisation, measuring its impact, and then two named frameworks: secure AI and responsible AI.

The adoption half is procedural. Recognise the types of gen AI solution, identify the factors that influence what an organisation needs, choose accordingly, work out the steps to integrate it, and identify techniques to measure the impact. There is a real discipline hiding in that last one, and it is the question most likely to reflect what a reader’s own employer is struggling with.

The security half names Google’s Secure AI Framework directly, alongside security across the machine learning lifecycle and the Google Cloud tools involved: secure-by-design infrastructure, Identity and Access Management, Security Command Center and workload monitoring. If you want the vendor-neutral counterpart that practitioners are also expected to recognise, the NIST AI risk management framework covers the same ground without a product attached.

Responsible AI closes the syllabus with transparency and privacy considerations. It is a small number of marks, it is easy to revise, and it is the part of the exam that most directly matches what a leader is actually asked in a board meeting.

What Are the GCP-GAIL Exam Format and Cost?

Fifty to sixty multiple choice questions in 90 minutes for $99, reported as a pass or fail result at approximately 70 percent, and scheduled through Google CertMetrics. That is the cheapest credential in Google’s certification range by a wide margin, and the only one that publishes its result as a verdict rather than a score.

Field Value
Exam code GCP-GAIL
Questions 50 to 60 multiple choice
Duration 90 minutes
Result Pass or fail, approximately 70 percent
Price $99 USD
Scheduling Google CertMetrics
Recommended training Generative AI Leader learning path and study guide

Ninety minutes for up to sixty questions is 90 seconds each, which is comfortable for a paper made of definitions and product matching. The pressure here is breadth, not time. If you finish early on this exam it is usually because you did not know something rather than because you knew it quickly.

The $99 price is the reason this credential spreads through organisations in groups. It is cheap enough for a manager to expense without a conversation, and Google publishes a free collection of preparation courses alongside it, which removes the other usual barrier. Details of what is and is not published sit on the official certification page.

Who Is This Credential Actually For?

Managers, strategists, product owners, analysts, consultants and anyone who has to make or defend a decision about generative AI without writing the code. Google says so directly, and the syllabus backs it up: there is no hands-on requirement, no prerequisite, and nothing that assumes prior cloud experience.

Engineers get something different from it. If you already build with these tools, GCP-GAIL is not a skills credential, it is a vocabulary and catalogue credential, and its value is in being able to talk to the business side without translating. That is worth 90 minutes for a lot of people, but be honest with yourself about which of the two you are buying.

Where it fits on a certification path

It sits outside the usual Google Cloud ladder rather than at the bottom of it. Nothing depends on it, and it does not lead anywhere in particular. If you want a technical Google credential in an adjacent area, the Google Cloud database engineer exam is a very different proposition and a useful contrast in what a hands-on Google paper looks like.

One published figure is worth knowing before you decide. Google’s own research, conducted with Ipsos across nine countries in late 2024, reports that more than 80 percent of people holding a Google Cloud certification say it opened doors to new opportunities and accelerated their path to promotion. That is a vendor’s own survey rather than an independent one, and it should be read as such, but it is at least a stated and sourced number rather than a vague claim.

How Should You Prepare for an Exam With No Console in It?

Treat it as a naming exercise and a matching exercise, in that order. There is no lab to build and no configuration to rehearse, so preparation is about turning a long list of terms and products into pairs you can recall under mild pressure. A week of short sessions beats a weekend of reading.

  1. Start by writing one-line definitions for the ten named fundamentals terms in your own words, paying particular attention to the difference between prompt tuning and prompt engineering, which the syllabus lists separately.
  2. Pair each stage of the machine learning lifecycle with the Google Cloud tool that serves it, because the syllabus asks for the pairing rather than for either list on its own.
  3. Group the Google offerings into four buckets, the platform, the ready-made assistants, the customer-facing suite and the API toolbox, and learn one concrete business use case for each bucket rather than for each product.
  4. Match every foundation model limitation to the practice that addresses it, so that hallucination pulls up grounding and retrieval-augmented generation, and a knowledge cutoff pulls up grounding with current data.
  5. Learn the prompting techniques as an ordered progression from zero-shot through few-shot to chain-of-thought and ReAct, and be able to say when each is worth the extra tokens.
  6. Finish with the sampling parameters and the two governance frameworks, since temperature, top-p, safety settings and output length are pure recall and the secure and responsible AI objectives are the easiest marks on the paper.

One useful reality check while you study: the wider industry picture is moving fast, and the developer survey on AI is a good vendor-neutral read on how these tools are actually being adopted, which makes the business-strategy domain feel less abstract.

Frequently Asked Questions

How many questions are on the GCP-GAIL exam?

Fifty to sixty multiple choice questions in 90 minutes. That works out at roughly 90 seconds per question, which is comfortable for a paper built from definitions and product matching.

What is the passing score for the Generative AI Leader certification?

The result is reported as pass or fail rather than as a numeric score, at approximately 70 percent. You will not receive a percentage, so there is no partial credit to plan around.

How much does the Generative AI Leader certification cost?

$99 USD, which makes it the least expensive credential in Google’s certification range. Google also publishes a free collection of preparation courses alongside it.

Do you need technical experience to take GCP-GAIL?

No. There is no hands-on requirement and no prerequisite. Google designed the credential for non-technical professionals such as managers, strategists and leaders, and every syllabus objective uses a verb like describe, identify or explain.

Which domain carries the most marks?

Google Cloud’s own gen AI offerings, at around 35 percent, or roughly 18 to 21 questions. It names about twenty five Google products and APIs, which is why grouping them by function beats memorising a flat list.

Which Google models does the exam name?

Gemini, Gemma, Imagen and Veo. You are expected to identify the use cases and strengths of each rather than to know how any of them was built.

What prompting techniques are on the syllabus?

Zero-shot, one-shot, few-shot, role prompting and prompt chaining, plus chain-of-thought and ReAct as the advanced pair. Each is named explicitly, so expect to match a technique to a use case.

What is the difference between grounding and fine-tuning here?

Grounding connects a model’s answers to a source of data at the time of the request, which the syllabus splits into first-party enterprise data, third-party data and world data. Fine-tuning changes the model itself. Both appear as remedies for foundation model limitations, and questions can turn on choosing between them.

How is the exam delivered and booked?

Through Google CertMetrics. The recommended preparation is Google’s own Generative AI Leader learning path and study guide, both named on the exam specification.

Is GCP-GAIL worth taking if you already work with these tools?

It depends what you want from it. It is a vocabulary and catalogue credential rather than a skills one, so its value for a practitioner is in being able to discuss options with the business side fluently, not in proving you can build anything.

Conclusion

GCP-GAIL is an unusual exam and a deliberately unusual one. Ninety minutes, 50 to 60 questions, $99, pass or fail at roughly 70 percent, and four domains that ask you to describe rather than to do. The largest of them is Google’s own product catalogue, and the smallest is the one about deciding whether any of it is worth adopting.

Prepare accordingly: define the terms precisely, pair the lifecycle stages with their tools, group the products by function, and match every model limitation to its remedy. If you want a second view of the exam’s shape before you commit, the GCP-GAIL study companion sets out the same ground from a practice-first angle. Then book it, because at $99 the cost of finding out is genuinely low.

Rating: 5 / 5 (1 votes)

The post Generative AI Leader Certification: Google’s AI Exam That Never Opens a Console appeared first on Big Data Prep.

]]>