01. In a retail pricing script, df is a Snowpark DataFrame over a table of several million products, and no action has been called on it. The developer writes:
marked_up = df["price"] * 1.2
What does marked_up hold after this line runs?
a) A list of Python floats, one for each price value that the line fetched from the table
b) A new DataFrame with a single computed column, which still needs an action to be evaluated
c) A pandas Series that was computed on the client when Python evaluated the * operator
d) A Column expression that becomes part of the SQL a later action runs
02. An IoT telemetry service keeps one Snowpark session open for days. Each processing step calls df.cache_result() on an expensive DataFrame, and the temporary tables behind those cached results pile up for as long as the session lives.
Which approaches remove a cached result's temporary table as soon as its step has finished, while the session stays open?
(Select two.)
a) Call drop_table() on the Table that cache_result() returned, once the step is done
b) Call delete() with no condition on the cached Table
c) Open it as with df.cache_result() as cached: and run the step inside the block
d) Call session.close() when the step ends, then create a new session for the step that follows
e) Call cache_result() again on the same DataFrame to replace it
03. The table shipments has three VARCHAR columns, in the order ORIGIN, DESTINATION, CARRIER. A logistics job builds new_rows with select("CARRIER", "ORIGIN", "DESTINATION") and runs:
new_rows.write.save_as_table("shipments", mode="append")
The call raises no error, but in the appended rows ORIGIN holds carrier names, DESTINATION holds origins and CARRIER holds destinations.
Which change loads each value into its matching column?
a) Pass column_order="index" to save_as_table, which pairs DataFrame columns with table columns by their names
b) Relabel the columns of new_rows as ORIGIN, DESTINATION, CARRIER with rename before the write, which puts them in the table's order
c) Use new_rows.write.insert_into("shipments"), which reorders columns to fit the table
d) Pass column_order="name" to save_as_table, which pairs columns by name instead of by sequence
04. An HR DataFrame df holds three nullable string columns: mobile_phone, home_phone and work_phone. A new column preferred_phone must hold the mobile number when there is one, else the home number, else the work number.
Which call produces that column?
a) df.with_column("preferred_phone", coalesce(col("mobile_phone"), col("home_phone"), col("work_phone")))
b) df.with_column("preferred_phone", nvl2(col("mobile_phone"), col("home_phone"), col("work_phone")))
c) df.with_column("preferred_phone", iff(col("mobile_phone").is_null() | col("home_phone").is_null(), col("work_phone"), col("mobile_phone")))
d) df.with_column("preferred_phone", col("mobile_phone")).fillna("home_phone", subset=["preferred_phone"])
05. A DataFrame readings holds sensor telemetry with the columns sensor, ts (a TIMESTAMP) and temp. Devices report at irregular intervals: sometimes every few seconds, sometimes once in twenty minutes. For every reading, an engineer needs the average temp of the readings from the same sensor taken within one hour before or after it, computed as avg(col("temp")).over(w).
Which definition of w produces that result?
a) Window.partition_by("sensor").order_by("ts").rows_between(-60, 60)
b) Window.partition_by("sensor").order_by("ts").rows_between(-make_interval(hours=1), make_interval(hours=1))
c) Window.partition_by("sensor").order_by("ts").range_between(-make_interval(hours=1), make_interval(hours=1))
d) Window.partition_by("sensor").order_by("ts").range_between(-3600, 3600)
06. During a review of an HR cleanup script, a developer finds these calls on t = session.table("employees"), where leavers is a DataFrame of employee ids:
t.update({"ACTIVE": False})
t.delete(source=leavers)
The author intended to deactivate, and then remove, just the employees listed in leavers.
Which statements about these calls are true?
(Select two.)
a) The update call is lazy, and with or without a condition no row changes until collect() is called on its result
b) The update call has no condition, so it sets ACTIVE to False on every row of the table and reports the count in rows_updated
c) The delete call joins the table to leavers on shared column names and removes the matched rows
d) The update call fails, as an assignment value must be a Column such as lit(False)
e) The delete call is invalid as written, as a condition must be provided when a source DataFrame is provided
07. A telemetry service uses a thread pool to run several independent Snowpark queries at once, and its design shares one thread-safe session among the worker threads. Two new requirements arrive: some workers must each run their own multi-statement transaction at the same time as the others, and one worker must query with a different warehouse and schema than the rest.
Which design decisions are sound?
(Select two.)
a) Have that one worker call session.use_warehouse() and session.use_schema() on the shared session before its query
b) Run the concurrent transactions on the shared session, with each worker calling begin_transaction() and commit() itself
c) Create separate session objects for the workers that run their own transactions and for the worker with a different warehouse and schema
d) Drop the shared session, as threads from a pool are unable to run queries on one session at the same time
e) Keep the independent queries on the shared session, which can run them concurrently
08. Sensor readings arrive as CSV files under @iot_stage/readings/. Each file has one header line, followed by rows that hold a device ID, a numeric reading and a timestamp. A developer has built reading_schema, a StructType whose StructField objects describe those three fields.
Which statement creates a DataFrame over the files whose columns carry the names and types declared in reading_schema, without treating the header line as data?
a) session.read.option("skip_header", 1).csv("@iot_stage/readings/", reading_schema)
b) session.read.schema(reading_schema).option("skip_header", 1).csv("@iot_stage/readings/")
c) session.read.option("skip_header", 1).csv("@iot_stage/readings/")
d) session.read.schema(reading_schema).parquet("@iot_stage/readings/")
09. A stored procedure imports a compiled library that is built only for the x86 CPU architecture. It runs on an existing Snowpark-optimized warehouse named train_wh, which was created without any memory or architecture setting. The team wants to keep the memory per node that the warehouse has today.
Which statement provides the x86 architecture and keeps that memory?
a) ALTER WAREHOUSE train_wh SET WAREHOUSE_TYPE = 'SNOWPARK-OPTIMIZED' RESOURCE_CONSTRAINT = 'MEMORY_1X_x86';
b) ALTER WAREHOUSE train_wh SET WAREHOUSE_TYPE = 'MEMORY_16X_x86';
c) ALTER WAREHOUSE train_wh SET RESOURCE_CONSTRAINT = 'MEMORY_16X';
d) ALTER WAREHOUSE train_wh SET RESOURCE_CONSTRAINT = 'MEMORY_16X_x86';
10. An HR analyst loads staff = session.table("staff"), whose columns are emp_id, mgr_id and full_name, and tries to pair each employee with a manager:
pairs = staff.join(staff, staff["mgr_id"] == staff["emp_id"])
The call raises SnowparkJoinException.
Which change is the documented way to make this self-join work?
a) Clone staff with copy.copy() and join staff to the clone
b) Build the condition with col("mgr_id") == col("emp_id") from the functions module
c) Pass lsuffix and rsuffix to the join, which gives the two sides different column names
d) Use staff.natural_join(staff), which pairs the shared column names automatically