# Failed (remote) cylc task fails but server fails to notice

**URL:** <https://cylc.discourse.group/t/failed-remote-cylc-task-fails-but-server-fails-to-notice/507>\
**Category:** Cylc Support\
**Created:** [July 22, 2022, 12:32am UTC](https://cylc.discourse.group/t/failed-remote-cylc-task-fails-but-server-fails-to-notice/507 "2022-07-22T00:32:20Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![Clint\_Seinen](https://yyz2.discourse-cdn.com/free1/user_avatar/cylc.discourse.group/clint_seinen/32/185_2.png) [@Clint\_Seinen](https://cylc.discourse.group/u/Clint_Seinen)\
**Post date:** [July 22, 2022, 12:32am UTC](https://cylc.discourse.group/t/failed-remote-cylc-task-fails-but-server-fails-to-notice/507/1 "2022-07-22T00:32:20Z")

</div>

As part of my workflow, I am submitting jobs (via `cylc play`) to a remote machine with a configuration like:

```auto
[[be-mach]]
    cylc path = /home/sci123/miniconda3/envs/cylc-8.0rc3/bin
    job runner = pbs
    install target = localhost
    global init-script = """
       export WORK_SPACE=/path/to/platform/specific/scratch
    """

```

from a host defined like:

```auto
    [[fe-mach]]
        cylc path = /home/sci123/miniconda3/envs/cylc-8.0rc3/bin
        job runner = pbs
        install target = localhost
        global init-script = """
             export WORK_SPACE=/path/to/platform/specific/scratch
        """

```

however, when the job fails on the `be-mach` platform, it takes a looong time before the failure is recognized on `fe-mach`, where I believe the server is running (I think its a server? Sorry if I’m not using the terminology correctly).

_Eventually_ the failure is picked up, and I’ve been digging around the documentation and source code, and it looks like I might be hitting default poll timing of 15 minutes because the default `zmq/TCP` communication is failing? When I run `cylc play` without `--debug`, in the `job.err` file, I see

```auto
CylcError: the workflow is no longer running at fe-machlogin.webaddress:PORTNUM
It has moved to fe-machlogin:PORTNUM

```

i.e. the `webaddress` portion is removed?

Running it with `--debug`, I get:

```auto
2022-07-21T23:36:50Z DEBUG - zmq:send {'command': 'graphql', 'args': {'request_string': '\nmutation (\n $wFlows: [WorkflowID]!,\n $taskJob: String!,\n $eventTime: String,\n $messages: [[String]]\n) {\n message (\n workflows: $wFlows,\n taskJob: $taskJob,\n eventTime: $eventTime,\n messages: $messages\n ) {\n result\n }\n}\n', 'variables': {'wFlows': ['test-dev/run1'], 'taskJob': '60000101T0000Z/model_run/01', 'eventTime': '2022-07-21T23:36:50Z', 'messages': [['CRITICAL', 'failed/ERR']]}}, 'meta': {'prog': 'message', 'host': 'cmpnode-999', 'comms_method': 'zmq'}}
Traceback (most recent call last):
  File "/home/sci123/miniconda3/envs/cylc-8.0rc3/bin/cylc", line 10, in <module>
    sys.exit(main())
  File "/home/sci123/miniconda3/envs/cylc-8.0rc3/lib/python3.9/site-packages/cylc/flow/scripts/cylc.py", line 675, in main
    execute_cmd(command, *cmd_args)
  File "/home/sci123/miniconda3/envs/cylc-8.0rc3/lib/python3.9/site-packages/cylc/flow/scripts/cylc.py", line 283, in execute_cmd
    entry_point.resolve()(*args)
  File "/home/sci123/miniconda3/envs/cylc-8.0rc3/lib/python3.9/site-packages/cylc/flow/terminal.py", line 226, in wrapper
    wrapped_function(*wrapped_args, **wrapped_kwargs)
  File "/home/sci123/miniconda3/envs/cylc-8.0rc3/lib/python3.9/site-packages/cylc/flow/scripts/message.py", line 173, in main
    record_messages(workflow_id, task_job, messages)
  File "/home/sci123/miniconda3/envs/cylc-8.0rc3/lib/python3.9/site-packages/cylc/flow/task_message.py", line 88, in record_messages
    send_messages(workflow, task_job, messages, event_time)
  File "/home/sci123/miniconda3/envs/cylc-8.0rc3/lib/python3.9/site-packages/cylc/flow/task_message.py", line 129, in send_messages
    pclient('graphql', mutation_kwargs)
  File "/home/sci123/miniconda3/envs/cylc-8.0rc3/lib/python3.9/site-packages/cylc/flow/network/client.py", line 253, in serial_request
    self.loop.run_until_complete(task)
  File "/home/sci123/miniconda3/envs/cylc-8.0rc3/lib/python3.9/asyncio/base_events.py", line 647, in run_until_complete
    return future.result()
  File "/home/sci123/miniconda3/envs/cylc-8.0rc3/lib/python3.9/site-packages/cylc/flow/network/client.py", line 198, in async_request
    self.timeout_handler()
  File "/home/sci123/miniconda3/envs/cylc-8.0rc3/lib/python3.9/site-packages/cylc/flow/network/client.py", line 317, in _timeout_handler
    raise CylcError(

```

where the error text is the same about it not running any more, but moved to a slightly different adddress.

It sounds like I could change it so the communication method is the polling by default, but it also sounds like that is inefficient? It looks to me like I need to do something with the hostnames in the platform configuration to get the zmq communication working, but I’m not certain

---

<div class="post-metadata">

**Author:** ![oliver.sanders](https://yyz2.discourse-cdn.com/free1/user_avatar/cylc.discourse.group/oliver.sanders/32/110_2.png) [@oliver.sanders](https://cylc.discourse.group/u/oliver.sanders)\
**Post date:** [July 22, 2022, 9:17am UTC](https://cylc.discourse.group/t/failed-remote-cylc-task-fails-but-server-fails-to-notice/507/2 "2022-07-22T09:17:44Z")

</div>

> [@Clint\_Seinen](#):
>
> ```auto
> CylcError: the workflow is no longer running at fe-machlogin.webaddress:PORTNUM
> It has moved to fe-machlogin:PORTNUM
> 
> ```

This message can occur when a ZMQ connection (the default communication method) times out. I think this has been caused by network configuration issues.

It’s really difficult debugging network issues from a distance. I suspect that hosts on the network are not able to connect to each other using their FQDNs (Fully Qualified Domain Names) and or the FQDN of a host seen from one place is not the same as seen from another? Here’s some information I hope will help in tracking down the issue…

#### FQDNs

Because Cylc is a distributed system it needs to be able to move around the network. This means each host needs to have a unique hostname and, where applicable, hosts need to be able to see other hosts using this unique hostname.

For safety, by default Cylc uses the FQDN (Fully Qualified Domain Name) of a host when connecting to it. I.E. Cylc is using the value of `hostname -f` (the longer form with the `webaddress` bit) rather than the short name. Long story short, each host should be contactable via it’s FQDN from itself and from any other hosts which Cylc needs to connect to it from.

You may find the network configuration needs some small changes to allow hosts to connect to other using the FQDN.

For situations where you’re unable to influence the network configuration of the job hosts, Cylc allows you to choose a different method of host identification. The options are:

- `name` (default) - Uses hostname (FQDN).
- `address` - Uses the IP address.
- `hardwired` (last resort) - Allows you to manually hardcode a name/address for each host.

The address option might be worth a try.

> **[Global Configuration — Cylc 8.6.0 documentation](https://cylc.github.io/cylc-doc/latest/html/reference/config/global.html#global.cylc%5Bscheduler%5D%5Bhost%20self-identification%5D)**

#### Communication Methods

> It sounds like I could change it so the communication method is the polling by default, but it also sounds like that is inefficient?

Cylc has three communication methods:

- ZMQ (a protocol built on TCP) - The preferred approach. Requires open TCP sockets.
- SSH - Fallback if opening TCP sockets between hosts is not permitted.
- Polling - Last resort if TCP/SSH are blocked, uses pull rather than push communication.

In polling mode, Cylc uses `ssh` to connect to the job host to inspect submitted and running tasks at a configured time interval. This is “pull” rather than “push” communication so updates will be recieved less frequently. These `ssh` connections will add a little to network load so it’s advised not to configure Cylc to poll too frequently.

If TCP or SSH are permitted these methods are preferable.

> **[Global Configuration — Cylc 8.6.0 documentation](https://cylc.github.io/cylc-doc/latest/html/reference/config/global.html#global.cylc%5Bplatforms%5D%5B%3Cplatform%20name%3E%5Dcommunication%20method)**

#### Platform Configs

> It looks to me like I need to do something with the hostnames in the platform configuration to get the zmq communication working, but I’m not certain

We’ve written up some example platform configs which might help:

> **[Platform Configuration — Cylc 8.6.0 documentation](https://cylc.github.io/cylc-doc/latest/html/reference/config/writing-platform-configs.html)**

If you do not specify `hosts` then Cylc uses the platform name I.E. the following two examples are equivalent:

```auto
[platforms]
    [[foo]]
        hosts = foo

[platforms]
    [[foo]]

```

> to a remote machine with a configuration like from a host defined like:

Note that `[platforms]` configure the job hosts. The “scheduler” host (the place where the workflow process runs) defaults to the host where `cylc play` is run but can be configured by `run hosts`:

> **[Global Configuration — Cylc 8.6.0 documentation](https://cylc.github.io/cylc-doc/latest/html/reference/config/global.html#global.cylc%5Bscheduler%5D%5Brun%20hosts%5D)**

Note that `install target = localhost` means that the job hosts shares the same `$HOME` filesystem as the scheduler host (where the workflow runs). If this is not the case choose an arbitrary name e.g. `install target = hpc-filesystem` (this tells Cylc that it needs to install the workflow files on this platform).

Hope this helps,  
Oliver

---

<div class="post-metadata">

**Author:** ![Clint\_Seinen](https://yyz2.discourse-cdn.com/free1/user_avatar/cylc.discourse.group/clint_seinen/32/185_2.png) [@Clint\_Seinen](https://cylc.discourse.group/u/Clint_Seinen)\
**Post date:** [July 22, 2022, 6:36pm UTC](https://cylc.discourse.group/t/failed-remote-cylc-task-fails-but-server-fails-to-notice/507/3 "2022-07-22T18:36:45Z")

</div>

Thank you very much for the comprehensive answer @oliver.sanders ! I think this should give me some things to track down to try to figure out the issue. I think @swartn was seeing similar issues on our old system, so its probably related.

RE: the `install target`, yeah we are still using the hack that is discussed [here](https://cylc.discourse.group/t/setting-up-a-platform-config-for-machines-that-share-a-home-directory/446) because all our machines share a home but not where the scratch spaces are, so we make links under `cylc-run`

---

<div class="post-metadata">

**Author:** ![Clint\_Seinen](https://yyz2.discourse-cdn.com/free1/user_avatar/cylc.discourse.group/clint_seinen/32/185_2.png) [@Clint\_Seinen](https://cylc.discourse.group/u/Clint_Seinen)\
**Post date:** [July 22, 2022, 10:17pm UTC](https://cylc.discourse.group/t/failed-remote-cylc-task-fails-but-server-fails-to-notice/507/4 "2022-07-22T22:17:50Z")

</div>

It looks like adding:

```auto
[scheduler]
    [[host self-identification]]
        method = address
        target = internal.domain.ca

```

to my `global.cylc` has potentially fixed this @oliver.sanders ! Thanks for the guidance.

However, I was wondering if you could enlighten me on a few things:

1. under the default settings, while on the job host, how does it attempt to determine that FQDN of the machine that ran `cylc play`? It seems like a chicken and an egg type of problem where it can only get the FQDN if it already has some name for it? Or does the job host used some stored name, and then say

2. Related to 1., for the `CylcError` I saw above, would it be the job host that first generated the `fe-machlogin.webaddress` name, or was it the server host? When I’m on the server host, and run `hostname -f`, I get the name that _doesn’t_ have the `webaddress`, but if I’m on the job host and run `ssh originname "hostname -f"` I get the same thing? So I’m not sure whos making the one with the `webaddress` portion.

3. When using the `address`/`target` combo, how does `cylc` use the given domain to figure it out?

---

<div class="post-metadata">

**Author:** ![hilary.j.oliver](https://yyz2.discourse-cdn.com/free1/user_avatar/cylc.discourse.group/hilary.j.oliver/32/4_2.png) [@hilary.j.oliver](https://cylc.discourse.group/u/hilary.j.oliver)\
**Post date:** [July 23, 2022, 3:52am UTC](https://cylc.discourse.group/t/failed-remote-cylc-task-fails-but-server-fails-to-notice/507/5 "2022-07-23T03:52:26Z")

</div>

> [@Clint\_Seinen](#):
>
> however, when the job fails on the `be-mach` platform, it takes a looong time before the failure is recognized on `fe-mach`, where I believe the server is running (I think its a server? Sorry if I’m not using the terminology correctly).

Presumably the same goes for job success, not just failure? (The same systems are used to communicate any task job completion or message).

> under the default settings, while on the job host, how does it attempt to determine that FQDN of the machine that ran `cylc play`?

The scheduler (started by `cylc play`) has to “self-identify” its location to its jobs, so that they know where to report back their status etc. It does this by putting a `contact` file on the job host, which contains the scheduler hostname or IP address, and port. When the Cylc job wrapper tries to report job status back to the scheduler it just uses the location in the contact file.

The scheduler uses `socket.getfqdn()` in Python get its own FQDN (which should match `hostname -f` in the shell). However, sometimes the FQDN returned on the scheduler host is not what’s needed to contact the scheduler host from other hosts on the network. That’s down to your network configuration. Which is why we provide other self-identification methods, including hardwiring if needed.

(I think this answers your second question too?)

> [@Clint\_Seinen](#):
>
> When using the `address`/`target` combo, how does `cylc` use the given domain to figure it out?

That’s done in `cylc/flow/hostuserutil.py`. According to the module docstring, we are using this method: [Get local IP Address with Python | Linux-Support.com](https://web.archive.org/web/20140606052543/http://www.linux-support.com/cms/get-local-ip-address-with-python/)

---

<div class="post-metadata">

**Author:** ![Clint\_Seinen](https://yyz2.discourse-cdn.com/free1/user_avatar/cylc.discourse.group/clint_seinen/32/185_2.png) [@Clint\_Seinen](https://cylc.discourse.group/u/Clint_Seinen)\
**Post date:** [July 28, 2022, 4:27pm UTC](https://cylc.discourse.group/t/failed-remote-cylc-task-fails-but-server-fails-to-notice/507/6 "2022-07-28T16:27:27Z")

</div>

Great thanks @hilary.j.oliver - that clears things up!

Regarding:

> Presumably the same goes for job success, not just failure? (The same systems are used to communicate any task job completion or message).

Yes, that is the case - I’m just early in the migration effort, so successes are rare 😉
