#6447 Upload outcome of the AtomicCI pipeline to amazon
Closed: Insufficient data by smooge. Opened by pingou.

I was asked today if the AtomicCI pipeline could upload its output to the amazon S3 cloud.

Since I have no idea if this is feasible or not, and if so how easy it would be, I'm asking here :)

It seems that both @bstinson and @dperpeet are interested to know this as well.

Thanks!


... and @dustymabe :)

please also see: https://pagure.io/fedora-infrastructure/issue/6022

please also see: https://pagure.io/fedora-infrastructure/issue/6022

Correct. But maybe it's easier to just upload these files for now, instead of switching over something that already exists. The artifacts (qcow2 images) don't need to be put anywhere else.

@bstinson do you have any idea of what you would need for this? (I assume credential, but do you know if you need anything else?)

@bstinson do you have any idea of what you would need for this? (I assume credential, but do you know if you need anything else?)

we (as fedora) probably would also need to determine a proper 'bucket' for the content to go into.

That being said I'd really like to do this all as part of #6022. There's really no reason we shouldn't be mirroring our other content in s3 as well and setting up the CDN to host it for everyone.

That being said I'd really like to do this all as part of #6022. There's really no reason we shouldn't be mirroring our other content in s3 as well and setting up the CDN to host it for everyone.

I most definitely not opposed to that, more wondering the time factor for one vs
the other.

Technically, getting credentials for putting things on AWS is trivial.
The questions we would need answered are: budget (please take that off-ticket and get approval from someone who'll pay), and what the main reason for this would be, given that the CentOS CI already has artifact storage.
Is it for end-user consumption, or something else?

Technically, getting credentials for putting things on AWS is trivial.
The questions we would need answered are: budget (please take that off-ticket and get approval from someone who'll pay),

This is 'free' according to #6022.

and what the main reason for this would be, given that the CentOS CI already has artifact storage.
Is it for end-user consumption, or something else?

Technically, getting credentials for putting things on AWS is trivial.
The questions we would need answered are: budget (please take that off-ticket and get approval from someone who'll pay),

This is 'free' according to #6022.

For that particular usecase, and up to a certain amount, yes.
That is why the usecase is important, and also how much this is going to be used.

The use case is publishing the developer stream as per https://pagure.io/fesco/issue/1774

If I'm not mistaken @bstinson, the current location in CentOS CI wouldn't deal well with too many consumers. I don't have any good numbers of what to expect regarding downloads.

Since we're looking at a stream, we wouldn't have to keep (m)any old images, I think.

I'm following up with @mattdm and others to find out what our limitations are (if any) from the budget side. Someone is working #6022 from the infra side already.

It would be great if the ticket were more specific about what material we want to store, and how (if at all) that affects tools or people that are expected to interact with that material. I'm not opposed to the concept at all, let's just be more concrete about what/why/how.

currently the artifacts are stored at http://artifacts.ci.centos.org/artifacts/fedora-atomic/f26/images/
As far as I know we only need those images and their logs.

currently the artifacts are stored at http://artifacts.ci.centos.org/artifacts/fedora-atomic/f26/images/
As far as I know we only need those images and their logs.

What about the actual ostree? http://artifacts.ci.centos.org/artifacts/fedora-atomic/f26/ostree/

So that answer the what part, but not the other two.

There is another question to add to Paul's list: for how long? Not for how long we want to use this but how long should images remain accessible there? (what's the retention policy?)

Also important is the "the current location in CentOS CI wouldn't deal well with too many consumers" remark: how many consumers would pull how much data how often?
That's also needed to estimate prices.

Yeah, if we can get some numbers, I'll run them by Amazon friends and see what they think. Thanks!

@dustymabe Do you have some numbers for the current Atomic Host? I think we can consider the developer stream to remain well below this.

I don't have exact numbers. We can get an approximation from @smooge, but I would anticipate the number of people pulling from this to be very low, mainly limited to core fedora contributors (not users) and even then they would only use it periodically.

@smooge have you had some time to look into these numbers?

I didn't even know anyone asked me for this... so the answer would be 'non'. I will see what I can get on this soon.

I have gone over the data and currently there are only 113 ips showing up for all of October with the majority of them only showing up once. Removing the internal boxes (10.5.x.x|RDU) from the list it drops down to 16 hosts which showed up more than one day and 84 which showed up once. The most requested version is 23 but that is from the RDU ip address. After that it is version 25.

We averaged 30 ips a day until the end of FLOCK and then we dropped down to 14 per day and then down to 6.5 last month. I had hoped it was a problem of collecting on my part but I have gone over all the logs and can't find any place the hosts could have logged to instead.

Soooo, I think for the purposes of this particular thing, that's good news. And if we get to the point where it's a problem, that will be different good news. :)

I have gone over the data and currently there are only 113 ips showing up for all of October with the majority of them only showing up once. Removing the internal boxes (10.5.x.x|RDU) from the list it drops down to 16 hosts which showed up more than one day and 84 which showed up once. The most requested version is 23 but that is from the RDU ip address. After that it is version 25.

I have a few questions:

  • What specific URLs are we checking? kojipkgs.fp.o or dl.fp.o?
  • We only do releases twice a month and we don't have any sort of auto update checker utility so maybe we could look at data for a longer period of time?
  • It is rather odd that the majority of machines would be hitting f23 or f25 at this point since we aren't doing releases for those any longer. We switched from dl.fp.o to kojipkgs.fp.o ~~about 90 days ago~~ in February for our updates. Could that be the reason we aren't seeing more of what we expect?

I have a few questions:

What specific URLs are we checking? kojipkgs.fp.o or dl.fp.o?
We only do releases twice a month and we don't have any sort of auto update checker utility so maybe we could look at data for a longer period of time?
It is rather odd that the majority of machines would be hitting f23 or f25 at this point since we aren't doing releases for those any longer. We switched from dl.fp.o to kojipkgs.fp.o about 90 days ago for our updates. Could that be the reason we aren't seeing more of what we expect?

While this is a very good point, even with some potentially missing data, are we expecting a different order of magnitude here? Otherwise I think it would be good to move forward with the Developer Stream. I think getting accurate information would be very helpful, I just want to be clear on whether we're blocking on this now or not. Thanks!

While this is a very good point, even with some potentially missing data, are we expecting a different order of magnitude here? Otherwise I think it would be good to move forward with the Developer Stream.

Yes. Moving forward here is desired. My questions were mainly from personal interest about the numbers.

The methodology was to look for a file that every system checks in for.

grep -h 'config.[[:space:]]200[[:space:]].ostree libsoup' $month/$day/$file where file was a list of systems that had shown up in previously. The majority of f23 pulls were from a Red Hat NAT ip so I don't know what systems are asking behind it. The f25 also come from that also. My guess is that some sort of CI is firing stuff up, testing etc and no one has turned it off when f23 EOLd. However I have no idea.

Good news everyone.. Dusty let me know that the libsoup had been replaced with a new user agent so my searches per day were wrong. This increased the usage amounts up to the same as they were before Flock. For September and October we saw a total of ~140 IP addresses show up over the month (removing the RH ips and the ci.centos.org one). There are an average of 30 checkins per day.

Can we move forward with this? We want to deliver this stream so that developers can benefit from a tested Atomic Host.

@mattdm and I discussed this (see also his earlier #comment-477805), and we agree based on what we've heard from AMZ it's OK to use our service (id 125523088429) for this stream. I'd like to see a specific retention policy, but that's not a blocker -- say, by the time we come back from DevConf.cz. @dperpeet, would you be able to work on that with relevant folks?

@pfrields Sure, I can. Since this is a stream that's supposed to be about the latest, I can't imagine why we would want to keep old images.

@dustymabe Do you have a good suggestion from the Atomic team on this?

My gut feeling is that we don't really need old images, but maybe we could keep the last 24 hours and then one each for 1, 3 and 7 days old or something.

@pfrields Sure, I can. Since this is a stream that's supposed to be about the latest, I can't imagine why we would want to keep old images.
@dustymabe Do you have a good suggestion from the Atomic team on this?
My gut feeling is that we don't really need old images, but maybe we could keep the last 24 hours and then one each for 1, 3 and 7 days old or something.

I think it depends on the users. So if I get a notification that my test fails on 01/02 and I don't check it until 01/10 (I was on holiday) and then I don't have time to look at the failure until 01/14, what do I do? Am I dead in the water? Has my test continued to run this whole time and fail on the latest build?

Basically is there anything I should need from the original 01/02 test failure or not? If no then keeping only a few days worth of stuff is ok. If I do need something from it then I'd prefer 14 days. Of course all of this depends on how much space each day accumulates.

So... what are the next steps here? Finishing 6022 and getting it into some amazon buckets?

Metadata Update from @kevin:
- Issue priority set to: Waiting on Reporter

Metadata Update from @smooge:
- Issue close_status updated to: Insufficient data
- Issue status updated to: Closed (was: Open)

Metadata