ancient-wolf-19325
09/04/2024, 1:26 PMyml at the end of this post.
What I'm observing seems a bit strange. When flyte first starts up it starts to schedule a lot of pods. I see up to 10K as pending but once the unprocessed queue depth goes down the system is scheduling hardly any work.
I would expect it to try to schedule 10s of thousands of pods with 70K workflows running. The roundtrip latency looks ok to my untrained eyes.
We checked our k8s API and can't find any throttling going on there - but maybe we are looking at the wrong things.
On the other hand I also see some tasks that are marked as Running for over 1 hour with the status message Sent to K8s... but no pod log link is available yet (see image below as well)
inline:
catalogcache:
reader:
maxItems: 100000
maxRetries: 3
workers: 100
writer:
maxItems: 100000
maxRetries: 3
workers: 100
workqueue:
workers: 1200
maxItems: 100000
config:
workers: 1200
maxItems: 100000
workqueue.Config:
workers: 1200
maxItems: 100000
workqueueConfig:
workers: 1200
maxItems: 100000
# Expose the metrics port for the Prometheus Service Monitor
propeller:
workers: 1800
prof-port: 10254
metrics-prefix: "flyte:"
max-workflow-retries: 50
kube-client-config:
qps: 4000 # Refers to max rate of requests (queries per second) to kube-apiserver
burst: 1200 # refers to max burst rate.
timeout: 30s # Refers to timeout when talking with the kube-apiserver
max-streak-lenght: 2
event:
rate: 10000
capacity: 20000
max-retries: 10ancient-wolf-19325
09/04/2024, 1:26 PM>># Answer It seems like you might be encountering issues related to Kubernetes API server throttling or FlytePropeller configuration. According to the Flyte performance documentation, even with increased
qps and burst values, the Kube API server can still throttle requests transparently, leading to performance drops. Monitoring specific metrics like workqueue_depth can help identify if throttling is the issue.
Additionally, the FlytePropeller configuration suggests that increasing the number of workers and adjusting the qps and burst values can help, but it requires careful tuning. The round_latency metric is crucial for understanding the performance of FlytePropeller.
Would you like more detailed steps on how to monitor and adjust these settings?
Sources:
- Flyte performance documentation
- FlytePropeller configuration
Slack Link
https://flyte-org.slack.com/archives/CP2HDHKE1/p1725456376.128309 👍 0 button 👎 0 button 💡 Hint: Mention @runllm in the thread for followups.