New RunLLM response generated: >>>*# Question* He...
# runllm-flyte-trial
a
New RunLLM response generated: >>>*# Question* Hello Team ! I recently implemented flyte on Hetzner cloud based kubernetes cluster built over k3s. I wanted to implement spark based tasks and can see clusters being initiated but all spark based clusters fail with following error log :
>>```++ id -u
+ myuid=1000 ++ id -g + mygid=1000 + set +e ++ getent passwd 1000 + uidentry=flytekit❌10001000/home/flytekit/bin/bash + set -e + '[' -z flytekit❌10001000/home/flytekit/bin/bash ']' + '[' -z /usr/local/openjdk-11 ']' + SPARK_CLASSPATH=':/opt/spark/jars/*' + env + grep SPARK_JAVA_OPT_ + sort -t_ -k4 -n + sed 's/[^=]*=\(.*\)/\1/g' + readarray -t SPARK_EXECUTOR_JAVA_OPTS + '[' -n '' ']' + '[' -z x ']' + export PYSPARK_PYTHON + '[' -z x ']' + export PYSPARK_DRIVER_PYTHON + '[' -n '' ']' + '[' -z ']' + '[' -z x ']' + SPARK_CLASSPATH='/opt/spark/conf::/opt/spark/jars/*' + case "$1" in + shift 1 + CMD=("$SPARK_HOME/bin/spark-submit" --conf "spark.driver.bindAddress=$SPARK_DRIVER_BIND_ADDRESS" --deploy-mode client "$@") + exec /usr/bin/tini -s -- /opt/spark/bin/spark-submit --conf spark.driver.bindAddress=10.244.3.55 --deploy-mode client --properties-file /opt/spark/conf/spark.properties --class org.apache.spark.deploy.PythonRunner local:///usr/local/bin/entrypoint.py pyflyte-fast-execute --additional-distribution s3://flyte/flytesnacks/development/AFUU2MUU5IGWXVCGNU4USDB6WU======/script_mode.tar.gz --dest-dir . -- pyflyte-execute --inputs s3://flyte/metadata/propeller/flytesnacks-development-asgp25xdskrs4btpqxmt/n0/data/inputs.pb --output-prefix s3://flyte/metadata/propeller/flytesnacks-development-asgp25xdskrs4btpqxmt/n0/data/0 --raw-output-data-prefix s3://flyte/data/eu/asgp25xdskrs4btpqxmt-n0-0 --checkpoint-path s3://flyte/data/eu/asgp25xdskrs4btpqxmt-n0-0/_flytecheckpoints --prev-checkpoint '""' --resolver flytekit.core.python_auto_container.default_task_resolver -- task-module hellospark task-name hello_spark 24/09/01 045937 WARN NativeCodeLoader: Unable to load native-hadoop library for your platform... using builtin-java classes where applicable Traceback (most recent call last): File "/usr/local/lib/python3.9/dist-packages/flytekit/core/data_persistence.py", line 560, in get_data self.get(remote_path, to_path=local_path, recursive=is_multipart, **kwargs) File "/usr/local/lib/python3.9/dist-packages/decorator.py", line 232, in fun return caller(func, *(extras + args), **kw) File "/usr/local/lib/python3.9/dist-packages/flytekit/core/data_persistence.py", line 122, in retry_request raise e File "/usr/local/lib/python3.9/dist-packages/flytekit/core/data_persistence.py", line 115, in retry_request return func(*args, **kwargs) File "/usr/local/lib/python3.9/dist-packages/flytekit/core/data_persistence.py", line 296, in get dst = file_system.get(from_path, to_path, recursive=recursive, **kwargs) File "/usr/local/lib/python3.9/dist-packages/fsspec/asyn.py", line 118, in wrapper return sync(self.loop, func, *args, **kwargs) File "/usr/local/lib/python3.9/dist-packages/fsspec/asyn.py", line 103, in sync raise return_result File "/usr/local/lib/python3.9/dist-packages/fsspec/asyn.py", line 56, in _runner result[0] = await coro ```
>>``` File "/usr/local/lib/python3.9/dist-packages/fsspec/asyn.py", line 634, in _get
rpaths = [ File "/usr/local/lib/python3.9/dist-packages/fsspec/asyn.py", line 635, in <listcomp> p for p in rpaths if not (trailing_sep(p) or await self._isdir(p)) File "/usr/local/lib/python3.9/dist-packages/s3fs/core.py", line 1483, in _isdir return bool(await self._lsdir(path)) File "/usr/local/lib/python3.9/dist-packages/s3fs/core.py", line 723, in _lsdir async for c in self._iterdir( File "/usr/local/lib/python3.9/dist-packages/s3fs/core.py", line 755, in _iterdir s3 = await self.get_s3(bucket) File "/usr/local/lib/python3.9/dist-packages/s3fs/core.py", line 353, in get_s3 return await self._s3creator.get_bucket_client(bucket) File "/usr/local/lib/python3.9/dist-packages/s3fs/utils.py", line 39, in get_bucket_client response = await general_client.head_bucket(Bucket=bucket_name) File "/usr/local/lib/python3.9/dist-packages/aiobotocore/client.py", line 394, in _make_api_call http, parsed_response = await self._make_request( File "/usr/local/lib/python3.9/dist-packages/aiobotocore/client.py", line 420, in _make_request return await self._endpoint.make_request( File "/usr/local/lib/python3.9/dist-packages/aiobotocore/endpoint.py", line 96, in _send_request request = await self.create_request(request_dict, operation_model) File "/usr/local/lib/python3.9/dist-packages/aiobotocore/endpoint.py", line 84, in create_request await self._event_emitter.emit( File "/usr/local/lib/python3.9/dist-packages/aiobotocore/hooks.py", line 66, in _emit response = await resolve_awaitable(handler(**kwargs)) File "/usr/local/lib/python3.9/dist-packages/aiobotocore/_helpers.py", line 6, in resolve_awaitable return await obj File "/usr/local/lib/python3.9/dist-packages/aiobotocore/signers.py", line 24, in handler return await self.sign(operation_name, request) File "/usr/local/lib/python3.9/dist-packages/aiobotocore/signers.py", line 90, in sign auth.add_auth(request) File "/usr/local/lib/python3.9/dist-packages/botocore/auth.py", line 423, in add_auth raise NoCredentialsError() botocore.exceptions.NoCredentialsError: Unable to locate credentials During handling of the above exception, another exception occurred: Traceback (most recent call last): File "/usr/local/bin/entrypoint.py", line 620, in <module> _pass_through() File "/usr/local/lib/python3.9/dist-packages/click/core.py", line 1157, in call return self.main(*args, **kwargs) File "/usr/local/lib/python3.9/dist-packages/click/core.py", line 1078, in main rv = self.invoke(ctx) File "/usr/local/lib/python3.9/dist-packages/click/core.py", line 1688, in invoke return _process_result(sub_ctx.command.invoke(sub_ctx)) File "/usr/local/lib/python3.9/dist-packages/click/core.py", line 1434, in invoke return ctx.invoke(self.callback, **ctx.params) ```
>>``` File "/usr/local/lib/python3.9/dist-packages/click/core.py", line 783, in invoke
return __callback(*args, **kwargs) File "/usr/local/bin/entrypoint.py", line 547, in fast_execute_task_cmd _download_distribution(additional_distribution, dest_dir) File "/usr/local/lib/python3.9/dist-packages/flytekit/core/utils.py", line 308, in wrapper return func(*args, **kwargs) File "/usr/local/lib/python3.9/dist-packages/flytekit/tools/fast_registration.py", line 151, in download_distribution FlyteContextManager.current_context().file_access.get_data( File "/usr/local/lib/python3.9/dist-packages/flytekit/core/data_persistence.py", line 564, in get_data raise FlyteAssertion( flytekit.exceptions.user.FlyteAssertion: USERAssertionError error=Failed to get data from s3://flyte/flytesnacks/development/AFUU2MUU5IGWXVCGNU4USDB6WU======/script_mode.tar.gz to ./ (recursive=False). Original exception: Unable to locate credentials 24/09/01 045940 INFO ShutdownHookManager: Shutdown hook called 24/09/01 045940 INFO ShutdownHookManager: Deleting directory /tmp/spark-7412c322-30de-459f-a770-ace11f1c698e```
>>
# Answer The error indicates that the Spark tasks are failing due to missing AWS credentials, which are required to access the S3 bucket. This is evident from the
NoCredentialsError
in the logs. 1. AWS Credentials Configuration: Ensure that the AWS credentials are correctly configured and accessible to the Spark tasks. This can be done by setting environment variables
AWS_ACCESS_KEY_ID
and
AWS_SECRET_ACCESS_KEY
in the Spark task running pod or container. 2. IAM Role Configuration: If you are using IAM roles, ensure that the role has the necessary permissions to access the S3 bucket and is correctly associated with the Spark task. This can be done by annotating the service account with the IAM role ARN. Would you like more detailed steps on how to configure AWS credentials or IAM roles for your Spark tasks? Sources: - Flyte Issue #2771 - Flyte Discussion on Spark Tasks - Flyte Discussion on AWS Credentials Slack Link https://flyte-org.slack.com/archives/CP2HDHKE1/p1725167390.951289 👍 0 button 👎 0 button 💡 Hint: Mention @runllm in the thread for followups.