2024-05-01 15:04:20 +00:00
|
|
|
This is a summary todo covering several subprojects, which would extend
|
|
|
|
git-annex to be able to use proxies which sit in front of a cluster of
|
|
|
|
repositories.
|
|
|
|
|
2024-05-01 16:19:12 +00:00
|
|
|
1. [[design/passthrough_proxy]]
|
2024-05-01 19:26:51 +00:00
|
|
|
2. [[design/p2p_protocol_over_http]]
|
|
|
|
3. [[design/balanced_preferred_content]]
|
|
|
|
4. [[todo/track_free_space_in_repos_via_git-annex_branch]]
|
|
|
|
5. [[todo/proving_preferred_content_behavior]]
|
2024-05-01 15:04:20 +00:00
|
|
|
|
2024-07-01 15:38:29 +00:00
|
|
|
## table of contents
|
|
|
|
|
|
|
|
[[!toc ]]
|
|
|
|
|
|
|
|
## planned schedule
|
|
|
|
|
2024-06-04 10:53:59 +00:00
|
|
|
Joey has received funding to work on this.
|
|
|
|
Planned schedule of work:
|
|
|
|
|
2024-06-27 19:28:10 +00:00
|
|
|
* June: git-annex proxies and clusters
|
2024-07-30 15:42:17 +00:00
|
|
|
* July: p2p protocol over http
|
|
|
|
* August, part 1: git-annex proxy support for exporttree
|
2024-08-08 20:06:02 +00:00
|
|
|
* August, part 2: [[track_free_space_in_repos_via_git-annex_branch]]
|
|
|
|
* September, part 1: balanced preferred content
|
2024-07-30 15:42:17 +00:00
|
|
|
* September, part 2: streaming through proxy to special remotes (especially S3)
|
|
|
|
* October, part 1: streaming through proxy continued
|
|
|
|
* October, part 2: proving behavior of balanced preferred content with proxies
|
2024-06-04 10:53:59 +00:00
|
|
|
|
2024-05-01 15:04:20 +00:00
|
|
|
[[!tag projects/openneuro]]
|
2024-06-04 11:51:33 +00:00
|
|
|
|
2024-07-01 15:38:29 +00:00
|
|
|
## work notes
|
2024-06-04 11:51:33 +00:00
|
|
|
|
2024-08-21 21:56:06 +00:00
|
|
|
* `git-annex assist --rebalance` of `balanced=foo:2`
|
|
|
|
sometimes needs several runs to stabalize.
|
|
|
|
|
2024-08-22 11:53:56 +00:00
|
|
|
May not be a bug, needs reproducing and analysis.
|
|
|
|
|
2024-08-17 17:30:24 +00:00
|
|
|
* Concurrency issues with RepoSizes calculation and balanced content:
|
2024-08-15 17:41:47 +00:00
|
|
|
|
|
|
|
* What if 2 concurrent threads are considering sending two different
|
|
|
|
keys to a repo at the same time. It can hold either but not both.
|
2024-08-17 17:30:24 +00:00
|
|
|
It should avoid sending both in this situation.
|
2024-08-15 17:41:47 +00:00
|
|
|
|
2024-08-15 17:50:50 +00:00
|
|
|
* There can also be a race with 2 concurrent threads where one just
|
|
|
|
finished sending to a repo, but has not yet updated the location log.
|
|
|
|
So the other one won't see an updated repo size.
|
|
|
|
|
|
|
|
The fact that location log changes happen in CommandCleanup makes
|
|
|
|
this difficult to fix.
|
|
|
|
|
|
|
|
Could provisionally update Annex.reposizes before starting to send a
|
|
|
|
key, and roll it back if the send fails. But then Logs.Location
|
|
|
|
would update Annex.reposizes redundantly. So would need to remember
|
|
|
|
the provisional update was made until that is called.... But what if it
|
|
|
|
is never called for some reason?
|
|
|
|
|
2024-08-15 20:15:48 +00:00
|
|
|
Also, in a race between two threads at the checking preferred content
|
|
|
|
stage, neither would have started sending yet, and so both would think
|
|
|
|
it was ok for them to.
|
|
|
|
|
2024-08-15 17:50:50 +00:00
|
|
|
This race only really matters when the repo becomes full,
|
|
|
|
then the second thread will fail to send because it's full. Or will
|
|
|
|
send more than the configured maxsize. Still this would be good to
|
|
|
|
fix.
|
|
|
|
|
2024-08-15 20:15:48 +00:00
|
|
|
* If all the above thread concurrency problems are fixed, separate
|
|
|
|
processes will still have concurrency problems. One case where that is
|
|
|
|
bad is a cluster accessed via ssh. Each connection to the cluster is
|
|
|
|
a separate process. So each will be unaware of changes made by others.
|
|
|
|
When `git-annex copy --to cluster -Jn` is used, this makes a single
|
|
|
|
command behave non-ideally, the same as the thread concurrency
|
|
|
|
problems.
|
|
|
|
|
2024-08-23 15:19:38 +00:00
|
|
|
* Possible solution:
|
|
|
|
|
|
|
|
Add to reposizes db a table for live updates.
|
|
|
|
Listing process ID, thread ID, UUID, key, addition or removal
|
2024-08-23 16:51:00 +00:00
|
|
|
(done)
|
2024-08-23 15:19:38 +00:00
|
|
|
|
|
|
|
Make checking the balanced preferred content limit record a
|
|
|
|
live update in the table and use other live updates in making its
|
|
|
|
decision. With locking as necessary.
|
|
|
|
|
|
|
|
Note: This will only work when preferred content is being checked.
|
|
|
|
If a git-annex copy without --auto is run, for example, it won't
|
|
|
|
tell other processes that it is in the process of filling up a remote.
|
|
|
|
That seems ok though, because if the user is running a command like
|
|
|
|
that, they are ok with a remote filling up.
|
|
|
|
|
|
|
|
In the unlikely event that one thread of a process is storing a key and
|
|
|
|
another thread is dropping the same key from the same uuid, at the same
|
|
|
|
time, reconcile somehow. How? Or is this perhaps something that cannot
|
|
|
|
happen?
|
|
|
|
|
|
|
|
Also keep an in-memory cache of the live updates being performed by
|
|
|
|
the current process. For use in location log update as follows..
|
|
|
|
|
|
|
|
Make updating location log for a key that is in the in-memory cache
|
|
|
|
of the live update table update the db, removing it from that table,
|
|
|
|
and updating the in-memory reposizes. This needs to have
|
|
|
|
locking to make sure redundant information is never visible:
|
|
|
|
Take lock, journal update, remove from live update table.
|
|
|
|
|
|
|
|
Somehow detect when an upload (or drop) fails, and remove from the live
|
|
|
|
update table and in-memory cache. How? Possibly have a thread that
|
2024-08-23 15:45:36 +00:00
|
|
|
waits on an empty MVar. Thread MVar through somehow to location log
|
|
|
|
update. (Seems this would need checking preferred content to return
|
|
|
|
the MVar? Or alternatively, the MVar could be passed into it, which
|
|
|
|
seems better..) Fill MVar on location log update. If MVar gets
|
2024-08-23 15:19:38 +00:00
|
|
|
GCed without being filled, the thread will get an exception and can
|
|
|
|
remove from table and cache then. This does rely on GC behavior, but if
|
|
|
|
the GC takes some time, it will just cause a failed upload to take
|
|
|
|
longer to get removed from the table and cache, which will just prevent
|
|
|
|
another upload of a different key from running immediately.
|
2024-08-23 15:45:36 +00:00
|
|
|
(Need to check if MVar GC behavior operates like this.
|
|
|
|
See https://stackoverflow.com/questions/10871303/killing-a-thread-when-mvar-is-garbage-collected )
|
2024-08-23 15:19:38 +00:00
|
|
|
|
|
|
|
Have a counter in the reposizes table that is updated on write. This
|
|
|
|
can be used to quickly determine if it has changed. On every check of
|
|
|
|
balanced preferred content, check the counter, and if it's been changed
|
|
|
|
by another process, re-run calcRepoSizes. This would be expensive, but
|
|
|
|
it would only happen when another process is running at the same time.
|
|
|
|
The counter could also be a per-UUID counter, so two processes
|
|
|
|
operating on different remotes would not have overhead.
|
|
|
|
|
2024-08-23 15:45:36 +00:00
|
|
|
When loading the live update table, check if PIDs in it are still
|
2024-08-23 15:19:38 +00:00
|
|
|
running (and are still git-annex), and if not, remove stale entries
|
|
|
|
from it, which can accumulate when processes are interrupted.
|
|
|
|
Note that it will be ok for the wrong git-annex process, running again
|
|
|
|
at a pid to keep a stale item in the live update table, because that
|
|
|
|
is unlikely and exponentially unlikely to happen repeatedly, so stale
|
|
|
|
information will only be used for a short time.
|
|
|
|
|
2024-08-23 15:45:36 +00:00
|
|
|
But then, how to check if a PID is git-annex or not? /proc of course,
|
|
|
|
but what about other OS's? Windows?
|
|
|
|
|
|
|
|
Perhaps stale entries can be found in a different way. Require the live
|
|
|
|
update table to be updated with a timestamp every 5 minutes. The thread
|
|
|
|
that waits on the MVar can do that, as long as the transfer is running. If
|
|
|
|
interrupted, it will become stale in 5 minutes, which is probably good
|
|
|
|
enough? Could do it every minute, depending on overhead. This could
|
|
|
|
also be done by just repeatedly touching a file named with the processes's
|
|
|
|
pid in it, to avoid sqlite overhead.
|
|
|
|
|
2024-08-24 13:37:24 +00:00
|
|
|
* Check for TODO XXX markers
|
|
|
|
|
2024-08-23 20:35:12 +00:00
|
|
|
* Check all uses of NoLiveUpdate to see if a live update can be started and
|
|
|
|
performed there. There is one in Annex.Cluster in particular that needs a
|
2024-08-24 13:37:24 +00:00
|
|
|
live update.
|
2024-08-23 20:35:12 +00:00
|
|
|
|
2024-08-24 13:37:24 +00:00
|
|
|
* The assistant is using NoLiveUpdate, but it should be posssible to plumb
|
|
|
|
a LiveUpdate through it from preferred content checking to location log
|
|
|
|
updating.
|
2024-08-23 20:35:12 +00:00
|
|
|
|
2024-08-17 18:58:36 +00:00
|
|
|
* `git-annex info` in the limitedcalc path in cachedAllRepoData
|
|
|
|
double-counts redundant information from the journal due to using
|
|
|
|
overLocationLogs. In the other path it does not, and this should be fixed
|
|
|
|
for consistency and correctness.
|
|
|
|
|
2024-08-09 18:16:09 +00:00
|
|
|
## completed items for August's work on balanced preferred content
|
|
|
|
|
|
|
|
* Balanced preferred content basic implementation, including --rebalance
|
|
|
|
option.
|
2024-08-17 17:30:24 +00:00
|
|
|
* Implemented [[track_free_space_in_repos_via_git-annex_branch]]
|
2024-08-22 11:53:56 +00:00
|
|
|
* `git-annex maxsize`
|
|
|
|
* annex.fullybalancedthreshhold
|
2024-08-06 18:18:30 +00:00
|
|
|
|
2024-08-08 19:34:36 +00:00
|
|
|
## completed items for August's work on git-annex proxy support for exporttre
|
2024-08-07 15:49:53 +00:00
|
|
|
|
|
|
|
* Special remotes configured with exporttree=yes annexobjects=yes
|
|
|
|
can store objects in .git/annex/objects, as well as an exported tree.
|
|
|
|
|
|
|
|
* Support proxying to special remotes configured with
|
|
|
|
exporttree=yes annexobjects=yes.
|
|
|
|
|
|
|
|
* post-retrieve: When proxying is enabled for an exporttree=yes
|
|
|
|
special remote and the configured remote.name.annex-tracking-branch
|
|
|
|
is received, the tree is exported to the special remote.
|
|
|
|
|
|
|
|
* When getting from a P2P HTTP remote, prompt for credentials when
|
|
|
|
required, instead of failing.
|
|
|
|
|
2024-08-07 16:27:24 +00:00
|
|
|
* Prevent `updateproxy` and `updatecluster` from adding
|
|
|
|
an exporttree=yes special remote that does not have
|
|
|
|
annexobjects=yes, to avoid foot shooting.
|
|
|
|
|
2024-08-08 18:05:05 +00:00
|
|
|
* Implement `git-annex export treeish --to=foo --from=bar`, which
|
|
|
|
gets from bar as needed to send to foo. Make post-retrieve use
|
|
|
|
`--to=r --from=r` to handle the multiple files case.
|
post-receive: use the exporttree=yes remote as a source
This handles cases where a single key is used by multiple files in the
exported tree. When using `git-annex push`, the key's content gets
stored in the annexobjects location, and then when the branch is pushed,
it gets renamed from the annexobjects location to the first exported
file. For subsequent exported files, a copy of the content needs to be
made. This causes it to download the key from the remote in order to
upload another copy to it.
This is not needed when using `git push` followed by `git-annex copy --to`
the proxied remote, because the received key is stored at all export
locations then.
Also, fixed handling of the synced branch push, it was exporting master
when synced/master was pushed.
Note that currently, the first push to the remote does not see that it
is able to get a key from it in order to upload it back. It displays
"(not available)". The second push is able to. Since git-annex push
pushes first the synced branch and then the branch, this does end up
with a full export being made, but it is not quite right.
2024-08-08 16:28:12 +00:00
|
|
|
|
2024-07-28 18:22:44 +00:00
|
|
|
## items deferred until later for p2p protocol over http
|
|
|
|
|
2024-07-28 21:19:27 +00:00
|
|
|
* `git-annex p2phttp` should support serving several repositories at the same
|
|
|
|
time (not as proxied remotes), so that eg, every git-annex repository
|
|
|
|
on a server can be served on the same port.
|
|
|
|
|
2024-07-29 13:11:27 +00:00
|
|
|
* Support proxying to git remotes that use annex+http urls. This needs a
|
|
|
|
translation from P2P protocol to servant-client to P2P protocol.
|
|
|
|
|
|
|
|
* Should be possible to use a git-remote-annex annex::$uuid url as
|
|
|
|
remote.foo.url with remote.foo.annexUrl using annex+http, and so
|
|
|
|
not need a separate web server to serve the git repository. Doesn't work
|
|
|
|
currently because git-remote-annex urls only support special remotes.
|
|
|
|
It would need a new form of git-remote-annex url, eg:
|
|
|
|
annex::$uuid?annex+http://example.com/git-annex/
|
2024-07-28 18:22:44 +00:00
|
|
|
|
|
|
|
* `git-annex p2phttp` could support systemd socket activation. This would
|
|
|
|
allow making a systemd unit that listens on port 80.
|
|
|
|
|
2024-07-04 19:18:06 +00:00
|
|
|
## completed items for July's work on p2p protocol over http
|
|
|
|
|
2024-07-24 16:19:53 +00:00
|
|
|
* HTTP P2P protocol design [[design/p2p_protocol_over_http]].
|
2024-07-04 19:18:06 +00:00
|
|
|
|
2024-07-22 23:50:08 +00:00
|
|
|
* addressed [[doc/todo/P2P_locking_connection_drop_safety]]
|
2024-07-05 19:34:58 +00:00
|
|
|
|
2024-07-11 15:50:44 +00:00
|
|
|
* implemented server and client for HTTP P2P protocol
|
|
|
|
|
2024-07-22 19:02:08 +00:00
|
|
|
* added git-annex p2phttp command to serve HTTP P2P protocol
|
|
|
|
|
2024-07-23 19:37:36 +00:00
|
|
|
* Make git-annex p2phttp support https.
|
|
|
|
|
2024-07-24 16:19:53 +00:00
|
|
|
* Allow using annex+http urls in remote.name.annexUrl
|
|
|
|
|
2024-07-26 14:24:23 +00:00
|
|
|
* Make http server support proxying.
|
|
|
|
|
2024-07-28 14:36:22 +00:00
|
|
|
* Make http server support serving a cluster.
|
|
|
|
|
2024-07-02 20:16:37 +00:00
|
|
|
## items deferred until later for [[design/passthrough_proxy]]
|
2024-06-23 13:28:18 +00:00
|
|
|
|
2024-07-01 15:29:04 +00:00
|
|
|
* Check annex.diskreserve when proxying for special remotes
|
|
|
|
to avoid the proxy's disk filling up with the temporary object file
|
|
|
|
cached there.
|
|
|
|
|
2024-06-28 19:32:00 +00:00
|
|
|
* Resuming an interrupted download from proxied special remote makes the proxy
|
|
|
|
re-download the whole content. It could instead keep some of the
|
|
|
|
object files around when the client does not send SUCCESS. This would
|
|
|
|
use more disk, but without streaming, proxying a special remote already
|
|
|
|
needs some disk. And it could minimize to eg, the last 2 or so.
|
2024-06-28 21:07:01 +00:00
|
|
|
The design doc has some more thoughts about this.
|
2024-06-18 16:07:01 +00:00
|
|
|
|
2024-06-28 19:32:00 +00:00
|
|
|
* Streaming download from proxied special remotes. See design.
|
2024-07-01 15:29:04 +00:00
|
|
|
(Planned for September)
|
2024-06-27 16:57:08 +00:00
|
|
|
|
2024-07-01 15:33:07 +00:00
|
|
|
* When an upload to a cluster is distributed to multiple special remotes,
|
|
|
|
a temporary file is written for each one, which may even happen in
|
|
|
|
parallel. This is a lot of extra work and may use excess disk space.
|
|
|
|
It should be possible to only write a single temp file.
|
|
|
|
(With streaming this won't be an issue.)
|
|
|
|
|
2024-06-27 17:40:09 +00:00
|
|
|
* Indirect uploads when proxying for special remote
|
|
|
|
(to be considered). See design.
|
|
|
|
|
2024-06-27 18:36:55 +00:00
|
|
|
* Getting a key from a cluster currently picks from amoung
|
|
|
|
the lowest cost remotes at random. This could be smarter,
|
|
|
|
eg prefer to avoid using remotes that are doing other transfers at the
|
|
|
|
same time.
|
|
|
|
|
2024-06-27 19:21:03 +00:00
|
|
|
* The cost of a proxied node that is accessed via an intermediate gateway
|
|
|
|
is currently the same as a node accessed via the cluster gateway.
|
|
|
|
To fix this, there needs to be some way to tell how many hops through
|
|
|
|
gateways it takes to get to a node. Currently the only way is to
|
|
|
|
guess based on number of dashes in the node name, which is not satisfying.
|
|
|
|
|
|
|
|
Even counting hops is not very satisfying, one cluster gateway could
|
|
|
|
be much more expensive to traverse than another one.
|
|
|
|
|
|
|
|
If seriously tackling this, it might be worth making enough information
|
|
|
|
available to use spanning tree protocol for routing inside clusters.
|
2024-06-25 21:50:22 +00:00
|
|
|
|
2024-06-12 15:55:18 +00:00
|
|
|
* Optimise proxy speed. See design for ideas.
|
2024-06-04 11:51:33 +00:00
|
|
|
|
2024-07-28 17:31:30 +00:00
|
|
|
* Speed: A proxy to a local git repository spawns git-annex-shell
|
|
|
|
to communicate with it. It would be more efficient to operate
|
|
|
|
directly on the Remote. Especially when transferring content to/from it.
|
|
|
|
But: When a cluster has several nodes that are local git repositories,
|
|
|
|
and is sending data to all of them, this would need an alternate
|
|
|
|
interface than `storeKey`, which supports streaming, of chunks
|
|
|
|
of a ByteString.
|
|
|
|
|
2024-06-12 15:55:18 +00:00
|
|
|
* Use `sendfile()` to avoid data copying overhead when
|
|
|
|
`receiveBytes` is being fed right into `sendBytes`.
|
2024-06-25 21:26:26 +00:00
|
|
|
Library to use:
|
|
|
|
<https://hackage.haskell.org/package/hsyscall-0.4/docs/System-Syscall.html>
|
2024-06-04 11:51:33 +00:00
|
|
|
|
2024-06-12 15:55:18 +00:00
|
|
|
* Support using a proxy when its url is a P2P address.
|
|
|
|
(Eg tor-annex remotes.)
|
2024-06-23 16:31:00 +00:00
|
|
|
|
2024-07-01 15:38:29 +00:00
|
|
|
## completed items for June's work on [[design/passthrough_proxy]]:
|
2024-06-23 20:38:01 +00:00
|
|
|
|
|
|
|
* UUID discovery via git-annex branch. Add a log file listing UUIDs
|
|
|
|
accessible via proxy UUIDs. It also will contain the names
|
|
|
|
of the remotes that the proxy is a proxy for,
|
|
|
|
from the perspective of the proxy. (done)
|
|
|
|
|
|
|
|
* Add `git-annex updateproxy` command (done)
|
|
|
|
|
|
|
|
* Remote instantiation for proxies. (done)
|
|
|
|
|
|
|
|
* Implement git-annex-shell proxying to git remotes. (done)
|
|
|
|
|
|
|
|
* Proxy should update location tracking information for proxied remotes,
|
|
|
|
so it is available to other users who sync with it. (done)
|
|
|
|
|
2024-06-27 19:28:10 +00:00
|
|
|
* Implement `git-annex initcluster` and `git-annex updatecluster` commands (done)
|
2024-06-23 20:38:01 +00:00
|
|
|
|
|
|
|
* Implement cluster UUID insertation on location log load, and removal
|
|
|
|
on location log store. (done)
|
|
|
|
|
|
|
|
* Omit cluster UUIDs when constructing drop proofs, since lockcontent will
|
|
|
|
always fail on a cluster. (done)
|
|
|
|
|
|
|
|
* Don't count cluster UUID as a copy in numcopies checking etc. (done)
|
|
|
|
|
|
|
|
* Tab complete proxied remotes and clusters in eg --from option. (done)
|
|
|
|
|
|
|
|
* Getting a key from a cluster should proxy from one of the nodes that has
|
|
|
|
it. (done)
|
|
|
|
|
|
|
|
* Implement upload with fanout to multiple cluster nodes and reporting back
|
|
|
|
additional UUIDs over P2P protocol. (done)
|
|
|
|
|
|
|
|
* Implement cluster drops, trying to remove from all nodes, and returning
|
|
|
|
which UUIDs it was dropped from. (done)
|
|
|
|
|
|
|
|
* `git-annex testremote` works against proxied remote and cluster. (done)
|
2024-06-25 14:06:28 +00:00
|
|
|
|
|
|
|
* Avoid `git-annex sync --content` etc from operating on cluster nodes by
|
|
|
|
default since syncing with a cluster implicitly syncs with its nodes. (done)
|
2024-06-25 15:35:41 +00:00
|
|
|
|
|
|
|
* On upload to cluster, send to nodes where its preferred content, and not
|
|
|
|
to other nodes. (done)
|
2024-06-25 18:52:47 +00:00
|
|
|
|
|
|
|
* Support annex.jobs for clusters. (done)
|
|
|
|
|
2024-06-26 16:56:16 +00:00
|
|
|
* Add `git-annex extendcluster` command and extend `git-annex updatecluster`
|
|
|
|
to support clusters with multiple gateways. (done)
|
|
|
|
|
|
|
|
* Support proxying for a remote that is proxied by another gateway of
|
|
|
|
a cluster. (done)
|
2024-06-27 16:20:22 +00:00
|
|
|
|
|
|
|
* Support distributed clusters: Make a proxy for a cluster repeat
|
|
|
|
protocol messages on to any remotes that have the same UUID as
|
|
|
|
the cluster. Needs extension to P2P protocol to avoid cycles.
|
|
|
|
(done)
|
2024-06-27 19:21:03 +00:00
|
|
|
|
|
|
|
* Proxied cluster nodes should have slightly higher cost than the cluster
|
|
|
|
gateway. (done)
|
2024-06-28 19:32:00 +00:00
|
|
|
|
|
|
|
* Basic support for proxying special remotes. (But not exporttree=yes ones
|
|
|
|
yet.) (done)
|
2024-07-01 15:29:04 +00:00
|
|
|
|
|
|
|
* Tab complete remotes in all relevant commands (done)
|
|
|
|
|
|
|
|
* Display cluster and proxy information in git-annex info (done)
|