Mostrando entradas con la etiqueta 4K. Mostrar todas las entradas
Mostrando entradas con la etiqueta 4K. Mostrar todas las entradas

viernes, 13 de mayo de 2016

Planning to offer 4K UHD service? A study of toots to optimize the selection of content for bandwidth allocation and design for 4K UHD infrastructures


Ultra High Definition television is many things: more pixels, more color, more contrast, and higher frame rates. Of these parameters, “more pixels” is much more mature commercially and Ultra HD 4k TVs are taking their place in peoples’ homes. Yet, Ultra HD content and service offerings are playing catch-up. We don’t yet have enough experience to know what good Ultra HD 4k content is nor do we know how much bandwidth to allocate to deliver great Ultra HD experiences to consumers. In article, we describe techniques and tools that could be used to validate the quality of uncompressed and compressed Ultra HD 4k content so that we can plan bandwidth
resources with confidence

We will describe the statistical methods we use to validate Ultra HD 4k content, and will present some of our results. We will also explore the impact of high-efficiency video coding (HEVC) compression on the statistics of Ultra HD 4k content. The data and analysis we present are intended to provide tools and data that could be used to optimize bandwidth allocation and design Ultra HD 4k service offerings.

INTRODUCTION


Only a decade ago, high definition HD was the big new thing. With it came new wider 16:9 aspect ratio flat screen TVs that made the living room stylish in a way that old CRTs couldn’t match. Consumers delighted in the new better television experience. Studios, programmers, cable, telco, and satellite video providers delivered a new golden-age of television. HD is now table stakes most places, and where that is not yet the case, it will be soon enough.

Yet now, before we hardly got used to HD, we are talking about Ultra HD (UHD) with at least four times as many pixels as HD. In addition to and along with UHD, we are getting a brand new wave of television viewing options. The Internet has become a rival of legacy managed television distribution pipes. Over-the-top (OTT) bandwidth is now often large enough to support 4k UHD exploration. New compression technologies such as HEVC are now available to make better use of video distribution channels. And the television itself is no longer confined to the home. Every tablet, notebook, PC, and smartphone now has a part time job as a TV screen; and more and more of those evolved-from-computer TVs have pixel density and resolution to rival dedicated TV displays.


Is all that resolution going to make a difference to consumers? If yes, what bandwidth will 4k UHD programming need? Those are two big questions our industry is exploring with respect to planning UHD services; yet they are not independent questions.



4k UHD is still new enough in the studios and post-production houses that 4k-capable cameras, lenses, image sensors, and downstream processing are still being optimized. Can we be sure yet that the optics and post processing are preserving every bit of “4k” detail? On the distribution side, could video compression and multi-bitrate adaptive streaming protocols change the amount of visual detail to an extent that it could conceivably turn “4k” quality into something more like “HD” or even less? If the 4k content we have available today for bandwidth and video quality testing does

not truly have a “4k”-level of detail, then we could go astray and plan for less bandwidth than we might need for future 4k UHD services. If the 4k content we have available today is truly “4k”, then we should also want to be sure that we do not over compress and turn 4k UHD into something less impressive.
Indeed, during our UHD 4k testing, we have found several candidate test sequences that appeared normal to the eye but turned out to have unusual properties when examined mathematically. Such content could lead to wrong conclusions when planning for UHD 4k bandwidth and services.
In this article, we present mathematical techniques to help answer the question “How 4k is it?” Our method examines 4k UHD video to see if it has a statistically expectable distribution of spatial detail as a function of 2-dimensional spatial frequency. The benchmark for our statistical expectations is drawn from numerous studies of the statistics of natural scenes.
Our main objective in writing this article is to describe methodology that might be useful in helping to decide which 4k UHD content should be included in the video test library intended to be used for bandwidth and video quality planning.




SOURCES OF 4K UHD CONTENT

There are many online places from which to obtain 4k content that could be considered for testing purposes. Industry-focused sources include the European Broadcasting Union (EBU), CableLabs, and blenderfoundation. Stock footage typically intended for promotional projects, but which might also be considered for testing purposes, is available online sites such as Shutterstock5, NYC B.Roll6, NatureFootage, and others that can be found by searching keywords such as “4k stock.” Video-sharing sites such as YouTube and Vimeo host compressed 4k UHD content that could be candidates for testing certain kinds of 4k UHD services.




CAMERA CONSIDERATIONS

The quality of 4k content depends on the quality of the camera, the particulars of the post processing such as filtering and compression, and the skill of the camera operator and crew. 4k-capable cameras available today range from consumer camcorders to cream-of-thecrop professional 4k-cameras that are used to create premium cinema and television content. Even some smartphones boast 4k cameras. Lens quality and image sensor size are key issues in 4k capture. Obviously, consumer and prosumer grade cameras should not be expected to have the top-quality lenses and image sensors found in high-end professional cameras. Yet, even in high-end cameras one needs to consider the interaction between the lens and image sensor. At this point in time, the image sensors found in many 4k cameras are larger than HD image sensorsAs a result, depth-of-field tends to be shallower. Background and foreground details that are out of the plane of focus can be softer than they are in HD. Depth-of-field can be increased by decreasing the aperture, but at the expense on less light which can result in noisier video because of sensor noise. Longer exposure times could improve the amount of light captured, but then motion blur could become an issue. All of these opto-electrical items are capable of producing 4k content that has less spatial detail in the subject matter and more noise than would otherwise be expected. More important, such content could lead to wrong conclusions about the amount of bandwidth that will

be needed to deliver great 4k experiences to consumers.



COMPRESSION CONSIDERATIONS

Video compression works mainly by strategically reducing the amount of spatial detail in video. Each compressed video frame is predicted from previously stored frames as much as possible. Whatever is unpredictable is packaged as a residual signal and sent to decoders, but not before the residual is further refined by being converted into a signal having less numerical precision through a technique called quantization. In the MPEG family of compression standards (MPEG-2, AVC/H.264, and HEVC/H.265), quantization has the effect of preferentially reducing the spatial details that are associated with higher spatial frequencies. In addition, AVC and HEVC employ spatial blurring filters that reduce the noticeability of spatial discontinuities between coding blocks. Both frequency-sensitive quantization and spatial blurring can reduce the kind of fine detail that 4k aims to show off. Without that “4k” detail in our test content our bandwidth predictions could be off when the next even-better generation of 4k camera arrives.



ADAPTIVE STREAMING CONSIDERATIONS

Over-the-top streaming services have taken the lead in delivering 4k content to consumers. Such services typically employ one of several adaptive streaming protocols that enable video to play smoothly even when the consumer’s bandwidth fluctuates significantly. Each adaptive-streaming video player senses its available bandwidth and requests a segment of programming that fits within its capabilities. If enough bandwidth is available, a 4k-capable video player would select lightly-compressed high bitrate segments having full 4k resolution (3840x2160). More restricted bandwidth could force selection of more aggressively compressed versions of the content though still at full 4k

resolution. Even more restricted bandwidth can force selection of aggressively compressed sub-4k resolution (for example, 1920x1080 or 960x540 etc.). The modulation of the resolution and the compression level mean that 4k adaptivestreaming services might not always be fully 4k. Thus, any test content derived from such a source could impact how test results should be interpreted.



SPATIAL FREQUENCY

An image is normally thought of as a 2-dimensional array of pixels with each pixel being represented by red, green, and blue values (RGB) or luma and 2 chrominance channels (for example, YUV or YCbCr). An image can also be represented as a 2-dimensional array of spatial-frequency components as illustrated in the figure. The visual pixel-based image and the spatial-frequency representation of the visual image are interchangeable mathematically. They have identical information, just organized differently.

Figure 1. Illustration the representation of an image in terms of spatial frequencies. The visual pixel-based image (A)
can be represented as a 2-dimensional array of complex numbers using Fourier transform techniques. The absolute value of the complex numbers is shown as a 2-dimensional magnitude spectrum (B) in which brighter areas correspond to larger magnitude values. (Note that the log of the magnitude spectrum is shown in B to aid visualization. The horizontal and vertical frequency axes are shown relative to the corresponding Nyquist frequency (±1).) The magnitudes of the main horizontal and vertical spatial frequency axes are shown in C. The main horizontal spatial frequency axis corresponds to zero vertical frequency (blue arrows in B), and the main vertical spatial frequency axis corresponds to zero horizontal spatial frequency (red arrows in B). (Note that the magnitude spectrum is mirror symmetric around the 0,0 point (center of B & D) along the main horizontal and vertical axes. Thus only the values from 0 to 1 (Nyquist frequency) are shown in C.) The dashed lines in C indicate the 1/spatial frequency

statistical expectation for natural scenes (the 1/spatial frequency appears as a line in the log plot). A contour map of the log of the magnitude spectrum is shown in D. Contour maps provide useful gestalts of the overall 2D magnitude spectrum. The data shown in B, C, and D were obtained by averaging the magnitude spectrum of individual frames over 250 frames (5 seconds) of the SVT CrowdRun 2160p50 test sequence. (All video processing and analysis discussed in this were performed using MATLAB10 and ffmpeg11.)

Spatial-frequency representations of images can further be represented by a magnitude component and a phase component. The magnitude component, called the magnitude spectrum, provides information on how much of the overall variation within the visual (pixel-based) image can be attributed to a particular spatial frequency. (Spatial frequency is 2-dimensional having horizontal and vertical parts.) The phase component, called the phase spectrum (not shown), provides information on how the various spatial frequencies interact to create the features and details we recognize in images. 

For the purposes of this study, we find it useful to use contour maps of the log of the magnitude spectra (as shown in Figure 1.) to create a gestalt, the 2-dimensional spatial frequency composition of images.


STATISTICS OF NATURAL SCENES

Images of natural scenes have an interesting statistical property: They have spatialfrequency magnitude spectra that tend to fall off with increasing spatial frequency in proportion to the inverse of spatial frequency. The magnitude spectra of individual images can vary significantly, but as an ensemble-average statistical expectation, it can be said that “the magnitude spectra of images of natural scenes fall off as one-overspatial-frequency.” This statement applies to both horizontal and vertical spatial frequencies. Examples of images adhering to this statistical expectation are shown in Figure 2.
Figure 2. Examples of adherence to the 1/f statistical expectation. These 8 image series created and analyzed by McCarthy & Owen13 illustrate the well-established statistical expectation that the magnitude spectra of images of natural scenes tend to be inversely proportional to spatial frequency. (Note that inverse proportionality appears as a straight line in the log plots shown.) The images shown in A, B, C, and D are representatives of the “Texture” (closeups of grass, sand, etc.), “Graffiti” (urban art), “Tilden Park” (woodland scenes), and “Urban” (San Francisco street scenes) image series, respectively. Other image series are “Fall” (colorful Fall foliage), “Damp” (puddles and environments where one might expect to find newts), “Winter” (snow scenes in New England), and “Garden” (UC Berkeley botanical garden).
Note that “natural-scene” images are not limited to pictures of grass and trees and the like. Any visually complex image of a 3-dimensional environment tends to have the one over-frequency  characteristic, though man-made environments tend to have stronger vertical and horizontal bias than unaltered landscape. The one-over-frequency characteristic can also be thought of as a signature of scale-invariance, which refers to the way in which small image details and large image details are distributed. Images of text and simple graphics do not tend to have one-over-frequency magnitude spectra.



A BENCHMARK FOR SPATIAL DETAIL

In this study, we leverage the one-over-frequency statistical expectation to see if itholds for 4k UHD content we have considered for use in our lab tests. Examples of 4k
(2160p50) video sequences that largely adhere to the one-over-frequency statistical
expectation are shown in Figure 3. Examples of candidate UHD 4k (3840x2160) videosthat violate the one-over-frequency statistical expectation are shown in Figure 4.
According to the results shown in Figures 3 & 4, the test sequences shown in Figure 3
remain viable candidates to be used in experiments to explore UHD bandwidth
planning; but we would exclude the test sequences represented in Figure 4 from our
UHD 4k test library.
Figure 3. UHD 4k test sequences that have statically expectable magnitude spectra. The SVT UHD 4k test sequences (3840x2160 at 50 frames per second) “CrowdRun”, “ParkJoy”, “DucksTakeOff”, “InToTrees”, and “OldTownCross” are shown left to right in columns. The corresponding visual (pixel-based) image, log of the magnitude spectrum averaged over 250 frames (5 seconds), main horizontal and vertical axis components, and contour map of the log average magnitude spectrum are shown top to bottom in each column. Note that each of the sequences can be welldescribed by the one-over-frequency statistical expectation (dashed lines in the plot on the third row from the top); though there are some subtle deviations from statistical expectation (see Figure 5). Note also that the contour maps provide concise distinguishing information about each image sequence.
Figure 4. Example of candidate UHD test sequences that DO NOT have statically expectable magnitude spectra. Some UHD 4k test sequences that we have obtained from various sources that appeared normal to the eye were found to have spatial magnitude spectra that were inconsistent with statistical expectations. Typical deviations from statistical expectations included: notch-like frequency distortions; excessive or diminished high or low frequency spatial detail (non-one-over-frequency behavior); and extraneous noise (see Figure 5).
The test sequences shown in Figure 3 are broadly in line with statistical expectationshowever, they do show subtle deviations as illustrated in Figure 5. These deviations are mainly the presence of isolated narrow-band noise-like distortions and mild loss of highfrequency high spatial detail. The sole reason we present Figure 5 is to illustrate a method of scrutinizing candidate UHD 4k test content to an extent not possible with the eye alone. (It should be noted that the SVT UHD 4k test sequences shown in Figures 3 and 5 were produced 10 years ago, long before the emergence of UHD 4k as a consumer service, and thus were on the cutting edge of UHD 4k research and development.)
Figure 5. A closer scrutiny of subtle deviations from the statically expectable magnitude spectrum. Although most UHD 4k candidate test content matches the one-over-frequency statistical expectation in general, some sequences do show subtle deviations. As illustrates for the SVT UHD 4k test sequences, these subtle deviations typically take to the form of extraneous noise that show up as isolated peaks and less-than-expected levels of high-frequency spatial detail (indicated by arrows pointing down). Note that the “DucksTakeOff” sequence meets statistical expectations particularly well.




EFFECTIVE RESOLUTION

A key feature of adaptive streaming protocols is the inclusion of reduced-resolution versions of content in order to provide uninterrupted video service even when a consumer’s available bandwidth is significantly curtailed. Although compressed at resolution less than full 4k resolution, the content seen by a viewer is upconverted to 4k resolution by either a set top box or the television display itself. In this way, the effective resolution is less than the displayed resolution. UHD 4k displays have such high resolution, and upconversion algorithms have become so good, that it is sometimes difficult to see by eye if a particular video is pristine full resolution or if some upconversion has occurred in the preparation of the contentFigure 6 illustrates a method of analyzing the effective resolution of “4k” (3184x2160) resolution test content more quantitatively than can be done by eye. It is well-known that a reduced effective resolution correlates to a loss of high-frequency spatial detailThis loss could, in principle, be evident by inspecting the main horizontal and vertical axes of the magnitude spectrum. We find that is not always the case. Modern rescaling algorithms are very sophisticated and the difference between lowered effective resolution and full resolution can be subtle. Instead, we find that the contour maps of the log of the magnitude spectrum are a much more sensitive indicator of effective resolution. A test for accepting candidate UHD 4k test content into a master library could be along the lines of determining the average radius of the outermost contour, accepting test content only when a certain radius threshold is exceeded.

Figure 6. An example of using contour maps of magnitude spectra to examine effective spatial resolution. Some UHD 4k candidate test content could have lower effective resolution than full UHD 4k (3840x2160). We simulate such a situation by downscaling and then upscaling back to 3840x2160 resolution using ffmpeg. From left to right, the downscaled resolution is: unaltered 3840x2160; 1920x1080; 960x540; and 480x270 as an extremum. Note that examination of only the main horizontal and vertical axes of the magnitude spectrum (middle row) reveals some differences, most notably some reduction in high-frequency spatial detail and shift in the narrow-band noise; but these details are too subtle to make confident decisions, particularly when the original full resolution content is unavailable for comparison. The contour maps of the log of the average magnitude spectrum provide more clear-cut evidence. The contour levels are that same for all columns. Thus the constriction of the contours towards the center indicated that the magnitude spectrum narrows (loses high-frequency spatial detail) thus quantifying the reduced effective resolution. This is, of course, expected. The significance of this figure is that is illustrates that the contour map method can be a sensitive measure of effective resolution of candidate test video.



EFFECT OF HEVC COMPRESSION

Video compression changes the amount of spatial detail in video, but the extent to which spatial detail is lost depends of the content itself and the aggressiveness of compression; i.e. the target bitrate. In Figure 7 we demonstrate that our method of evaluating test content provides a way of testing the effective resolution of HEVC compressed content. Note that our method indicates that effective resolution is more sensitive to compression for some kinds of content compared to other kinds of content. As such, our contour-map method could be used to optimize the selection of compressed UHD 4k content for testing purposes in terms of both the intrinsic image characteristics of content and the impact of bit rate. In this way, our contour-map method could serve as a content-independent method of measuring effective resolution and thus selection of useable UHD 4k test content.
Figure 7. An example of using contour maps of magnitude spectra to examine the effect of video compression. Shown here are the contour maps of the average (250 frames, 5 seconds) of the log magnitude spectrum of each of the SVT UHD 4k test sequences compressed with HEVC to various extents. We used the libx26514 library with ffmpeg to perform the HEVC compression. The crf value noted at the top of each column indicates the value of the constant rate factor (crf) parameter used in the ffmpeg libx265 command line. Smaller values of crf created more lightly compressed video. Video compressed with a crf value of 50 is typically very heavily artifacted. Video compressed with a crf value of 10 produces contour maps that are very similar to those for uncompressed video (see Figure 3). Note that the impact of the crf value is content dependent. For “CrowdRun” and “ParkJoy”, the crf values below ~30 do not have a major impact on effective resolution. On the other hand, a noticeable change in effective resolution is evident for a crf value of 20 for “IntoToTrees”, “DucksTakeOff”, and “OldTownCross”. (The dashed lines provide a reference for the radial extent of outer contour of the lightly compressed and uncompressed versions of the video.)



DISCUSSION & CONCLUSIONS

The objective of this paper was to present techniques that might be useful in evaluating UHD 4k video sequences that could be candidates for testing related to planning UHD 4k products and services. Selection of test content that is not representative of anticipated UHD 4k programming – including future UHD 4k programming that will be available when the end-to-end UHD 4k ecosystem has been optimized – could lead to wrong conclusions about what bandwidth and level of video quality would be needed. We show in this study that ensemble-average statistical expectations related to spatial frequency magnitude spectra of images of natural scenes can be used as a benchmark of comparison to address the question: “How 4k is it?” We find that major deviations from statistical expectations can be considered grounds for excluding content from a test library. (Though perhaps some synthetic compression busters should be retained to stress compression equipment and distribution services.) We also find that examination of contour maps of the log of the magnitude spectra are sensitive indicators of effective resolution. Contour maps that indicate a lack of 4k-level effective resolution can be considered grounds for excluding content from a test libraryThe main horizontal and vertical components of the magnitude spectrum seem to be
good probes for detecting added noise and gross distortions; but they are not highly sensitive probes of effective resolution. 
Significantly, our contour-map method is also a sensitive content-independent probe that can be used to evaluate compressed content for inclusion in UHD 4k video test libraries to be used for planning UHD 4k video quality and bandwidth.



Test Videos: 4K (UHD), Bitrate tests, sample clips

4K UHD Test Clips


  1. 4K Test Patterns (in MP4 @ 30fps) (YouTube) (thanks hansolo)
  2. 23.976fps (in MP4) (YouTube)
  3. 24fps (in MP4) (YouTube)
  4. 25fps (in MP4) (YouTube)
  5. 29.970fps, 51Mbps (hdmkv's iPhone 6S 4K clip)
  6. 59.940fps (in MKV)
  7. 60fps (in MP4) (YouTube)
  8. H264, up to 30fps (thanks hansolo for #8-12)
  9. H264, 50-60fps
  10. H265 8bit, up to 30fps
  11. H265 10bit, up to 30fps
  12. H265 10bit, 50-60fps
  13. HDR 10bit HEVC, 24fps
  14. HDR 10bit HEVC, 59.940fps
  15. VP9 (open-source alternative to HEVC) (YouTube)
  16. 5K (5120x2700) (in MP4 @ 60fps) (YouTube)
  17. 8K (7680x4320) (in MP4 @ 29.970fps) (YouTube)

Bitrate Test Clips

Media players should be able play 70Mbps or better smoothly, w/o stutters for full 1080p, and 108-128Mbps for full 4K

Resources for Additional Test Clips or Samples


HD Audio Test Clips

  1. Dolby Digital Plus 5.1 (in M2TS @ 1080p/29.970) (thanks wesk05)
  2. Dolby Digital Plus 7.1 Channel Check (in MKV @ 1080p/29.970)
  3. Dolby TrueHD 5.1 (use clip #10 in 'Codecs' section below)
  4. Dolby TrueHD 7.1 Channel Check (in MKV @ 1080p/29.970)
  5. Dolby ATMOS 'Amaze' Demo (in M2TS @ 1080p/24.000)
  6. Dolby ATMOS '747' Audio Demo (in M2TS @ 1080p/29.970 with DD+ 5.1 secondary track) (thanks movie78)
  7. Dolby ATMOS 'Helicopter' Audio Demo (in M2TS @ 1080p/29.970 with DD+ 5.1 secondary track) (thanks movie78)
  8. DTS-HD HRA 5.1 (in MKV @ 1080p/23.976 VC-1)
  9. DTS-HD HRA 7.1 (in MKV @ 1080p/29.970)
  10. DTS-HD MA 5.1 Channel Check (in M2TS @ 1080p/23.976)
  11. DTS-HD MA 7.1 Speaker Phase (in MKV @ 1080p/23.976)
  12. DTS-HD MA 7.1 'Dredd' Audio Channel Check (in M2TS @ 1080p/23.976) (thanks looun)
  13. DTS:X 'All Around Us' Demo (in MKV @ 1080p/23.976) (thanks wesk05)
  14. DTS:X 'Gravity' Demo (in M2TS @ 1080p/23.976) (thanks wesk05)
  15. DTS:X 'Movement' Demo (in M2TS @ 1080p/23.976) (thanks wesk05)
  16. LPCM 5.1 (in MKV @ 1080p/23.976) (thanks wesk05)
  17. LPCM 7.1 (in MKV @ 1080p/23.976)
  18. AAC 5.1
  19. FLAC 5.1 (use clip #5 in '3D Test Clips' section above)

Library samples

A zipped collection of 1,000 empty movie files, with NFO files, poster, and fanart for each entry. Various movies from different years, including sequels/sets, remakes, movies named the same but unrelated, various genres, and so on. Useful for testing things like library scanning speed, library navigation, filtering, etc.

HEVC / H265 Verification test plan

ISO/IEC JTC1/SC29/WG11 N14226

Introduction

This document contains the plan for the video verification test to be conducted to verify the coding performance of the HEVC Main and Main 10 profiles. A formal subjective evaluation will be conducted comparing the HEVC Main and Main 10 profiles to the AVC High and High 10 profiles, respectively. A range of video resolutions from 480p to 4K will be tested.

Test conditions

The following test conditions will be used for the HEVC verification test.
1.      Number of sequences and video resolutions:
a.       5 sequences for each resolution (480p, 720p, 1080p and 4K)
2.      Bitstreams
a.       Generated with HM 12.1 for HEVC bitstreams
b.      Generated with JM 18.5 for AVC bitstreams
c.       In addition to a. and b., other HEVC and/or AVC bitstreams generated with encoders that are optimized for subjective quality may be tested if available.
3.      Encoding parameters
a.       Fixed QP.
4 bit rate points per sequences covering the whole MOS range as much as possible

b.      Bit depth
                                                              i.      8 bits for 480p, 720p and 1080p
                                                            ii.      8 and 10 bits for 4K

c.       Coding structure depending on the nature of the sequence.

i. Random access, RA (Storage/Streaming)
   - Intra refresh at approximately 1 second intervals.
   - Picture reordering allowed.

ii. Low delay, LD (Video conferencing)
    - No Intra refresh
    - Without picture reordering.


d.      Other settings as in the configuration files
                                                             
i. cfg/encoder_randomaccess_main.cfg, encoder_randomaccess_main10.cfg or encoder_lowdelay_main.cfg for HM

ii. bin/HM-like/encoder_JM_RA_B_HE.cfg or bin/HM-like/encoder_JM_LB_HE.cfg configurations for JM18.5

Test Sequences

The following test sequences are selected for the subjective test.
Table 1: Selected test sequences and properties

Sequence
Source
[Copyright]
Width x Height
Frame rate
Bit depth
Length (frames)
RA / LD
BT709Birthday
Technicolor [C3]
3840x2160
50
10
500
RA
Book
BBC [C4]
3840x2160
50
10
500
RA
manage
4EVER [C2]
3840x2160
60
8
600
RA
HomelessSleeping
Kamerawerk [C8]
3840x2160
60
8
600
RA
traffic
Plannet, Inc [C1]
4096x2048
30
8
300
RA
JohnnyLobby
Vidyo [C7]
1920x1080
60
8
600
LD
Calendar
BBC [C4]
1920x1080
50
8
500
RA
SVT15
SVT [C6]
1920x1080
50
8
500
RA
sedofCropped
4EVER [C2]
1920x1080
60
8
600
RA
UnderBoat1
NTIA [C5]
1920x1080
30
8
300
RA
ThreePeople
Vidyo [C7]
1280x720
60
8
600
LD
QuarterBackSneak1
NTIA [C5]
1280x720
30
8
300
RA
BT709Parakeets
Technicolor [C3]
1280x720
50
8
500
RA
SVT01a
SVT [C6]
1280x720
50
8
500
RA
SVT04a
SVT [C6]
1280x720
50
8
500
RA
Cubicle
Vidyo [C7]
832x480
30
8
300
LD
Anemone
NTIA [C5]
832x480
30
8
300
RA
BT709BirthdayFlash
Technicolor [C3]
832x480
50
8
500
RA
Ducks
Plannet, Inc [C1]
832x480
60
8
600
RA
WheelAndCalendar
BBC [C4]
832x480
50
8
500
RA

 Encoding Results

The following table shows the JM18.5 and HM12.1 encoding results on the sequences shown in Table 1. The QP parameters were selected such that the bitrate of the HM12.1 bitstreams are approximately half of the bitrate of the corresponding JM18.5 bitstreams.  The range of the QP values was also selected so that the subjective quality of the encoded sequence span as large a range of the MOS range as possible.

Table 2: JM18.5 and HM11.0 encoding results

JM18.5
HM12.1

QPISlice
kbps (a)
QPISlice
kbps (b)
Bitrate Difference (a - b)/b
4K
BT709Birthday
24
14064
27
6838
51%


30
7223
32
3654
49%


35
4461
37
2154
52%


40
2858
42
1317
54%

Book
22
10911
24
5742
47%


27
5805
29
2738
53%


32
3525
33
1643
53%


37
2214
37
1042
53%

HomelessSleeping
23
38876
25
16608
57%


26
12168
27
5526
55%


31
5617
31
2581
54%


37
3112
35
1488
52%

menage
27
36607
31
17840
51%

31
21261
35
10466
51%

35
12731
39
6139
52%

38
8819
42
4021
54%

traffic
27
13309
31
6205
53%

32
6583
36
3137
52%

37
3618
40
1844
49%

42
2090
44
1056
49%
1080p
JohnnyLobby
23
2761
24
1477
46%

(low delay)
27
895
28
445
50%

31
468
32
227
51%

35
298
36
139
54%

Calendar
23
3057
26
1407
54%

27
1668
30
787
53%

32
958
34
487
49%

36
686
38
322
53%

SVT15
28
6805
31
3549
48%

32
3767
35
1903
49%

36
2214
39
1028
54%

41
1109
43
547
51%

sedofCropped
27
13762
31
6345
54%


31
6726
35
3165
53%


35
3462
39
1619
53%


39
1863
42
971
48%

UnderBoat1
24
4407
27
1910
57%

29
2026
31
990
51%

33
1196
35
554
54%


37
729
39
325
55%
720p
ThreePeople
25
1414
28
648
54%

(low delay)
29
739
32
346
53%

33
433
36
200
54%

38
240
40
117
51%

BT709Parakeets
26
1151
30
553
52%

30
709
34
333
53%

33
499
37
232
54%

37
332
40
161
52%

QuarterBackSneak
22
3844
25
1959
49%


27
2039
30
1009
51%


32
1145
35
541
53%


37
694
39
336
52%

SVT01a
27
1271
31
594
53%


31
733
35
336
54%


35
435
38
215
51%


39
283
41
132
53%

SVT04a
28
4178
32
2154
48%

31
2665
35
1361
49%

34
1699
38
849
50%

37
1072
41
503
53%
480p
Cubicle
22
1014
24
505
50%

(low delay)
25
502
27
264
48%

30
210
32
106
49%

35
105
37
49
53%

Anemone
25
990
29
478
52%


29
581
33
271
53%


33
353
37
164
54%


38
202
41
99
51%

BT709BirthdayFlash
29
1515
33
774
49%


32
1003
37
499
50%


35
679
41
315
54%


39
399
44
205
49%

Ducks
27
2178
31
1033
53%


32
1063
36
517
51%


35
719
39
344
52%


38
492
42
226
54%

WheelAndCalender
22
1129
25
519
54%


27
513
30
243
53%


32
295
35
136
54%


37
190
39
93
51%

Description of testing environment and methodology


The test procedure foreseen for the formal subjective evaluation will consider two main requirements:
  • to be as much as possible reliable and effective in verifying the performance in terms of subjective quality (and therefore adhering the existing recommendations);
  • to take into account the evolution of technology and laboratory set-up oriented to the adoption of FPD (Flat Panel Display) and video server as video recording and playing equipment.
Therefore, one of the test methods described in [1] are planned to be used, applying some modification to them, in relation to the kind of display, the video recording and play-back equipment.

Test method

The test method adopted for this evaluation is DCR (Degradation Category Rating) [1].


A.1.1 Degradation Category Rating (DCR)

This test method is commonly adopted when the material to be evaluated shows a range of visual quality that well distributes across all quality scales.
This method will be used under the schema of evaluation of the quality (and not of the impairment); for this reason a quality rating scale made of 11 levels will be adopted, ranging from "0" (lowest quality) to "10" (highest quality). The test will be held in three different laboratories located in countries speaking different languages: This implies that it is better not to use categorical adjectives (e.g. excellent good fair etc.) to avoid any bias due to a possible different interpretation by naive subjects speaking different languages.
All the video material used for these tests will consist of video clips of 10 seconds duration.
The structure of the Basic Test Cell (BTC) of DCR method is made by two consecutive presentations of the video clip under test; at first the original version of the video clip is displayed, immediately afterwards the coded version of the video clip is presented; then a message displays for 5 seconds asking the viewers to vote (see Figure 1)


A.2 How to express the visual quality opinion with DCR

The viewers will be asked to express their vote putting a mark on a scoring sheet.
The scoring sheet for a DCR test is made of a section for each BTC; each section has a box wherein which the viewer shall write the score ranging from 0 to 10 (see Figure 2). By writing a score of “10”, the subject will express an opinion of “best” quality, while by writing a score of “0” the subject will express an opinion of “worst” quality.

The vote has to be written when the message "Vote N" appears on the screen. The number "N" is a numerical progressive indication on the screen aiming to help the viewing subjects to use the appropriate box of the scoring sheet.

A.4 Training and stabilization phase

The outcome of a test is highly dependent on a proper training of the test subjects.
For this purpose, each subject has to be trained by means of a short practice (training) session.
The video material used for the training session must be different from those of the test, but the impairments introduced by the coding have to be as much as possible similar to those in the test.
The stabilization phase uses the test material of a test session; three BTCs, containing one sample of best quality, one of the worst quality and one of medium quality, are duplicated at the beginning of the test session. By this way, the test subjects have an immediate impression of the quality range they are expected to evaluate during that session.
The scores of the stabilization phase are discarded. Consistency of the behaviour of the subjects will be checked inserting in the session a BTC in which original is compared to original.

A.5 The laboratory set-up

The laboratory for a subjective assessment will be set up according to [1], except for the selection of the display and the video play-out server.

For 4K video clips, high quality LCD monitors will be used with diagonal size equal to or higher than 56'' and able to accept resolutions of up to 3840x2160. Play-out of 3840x2048 video clips is done at the native resolution using the central area of the screen; the remaining part of the screen is set to a mid grey level (128 in 0-255 range)". In the case where the width of the sequence exceeds 3840, the left and right sides of the picture would be cropped and only the centre 3840 pixels are shown.
For other resolutions, High quality LCD monitors (or TV set) are used, having a diagonal size equal or higher of 40” and capable to accept resolution equal to 1920 x 1080. When using TV sets all the local colour and contrast features must be disabled (where applicable).
Play-out of 1080p, 720p and 480p video clips is done at the native resolution using the central area of the screen; the remaining part of the screen is set to a mid grey level (128 in 0-255 range).
The video play server, or the PC, used to play video has to be able to support the display of 4K, 1080p, 720p and 480p video formats, at 24, 30, 50 and 60 frames per second, without any limitation, or without introducing any additional temporal or visual degradation.

A.5.1 Viewing distance

The viewing distance varies according to the physical dimensions of the active part of the video; this will lead to a viewing distance varying from 1.5H to 4H, where H is equal to the height of the active part of the screen, depending on the size of the active part of the screen and its native resolution.
The number of subjects seating in front of the monitor is a function of the monitor size and of the selected viewing distance.

A.5.2 Viewing environment.
The test laboratory has to be carefully protected from any external visual or audio pollution.
Internal general light has to be low (just enough to allow the viewing subjects to fill out the scoring sheets) and a uniform light has to be placed behind the monitor, in a way no direct light hits the viewing subjects seated in front of the screen; the light behind he monitor must be dimmed to an intensity as specified in Table 4 of Recommendation ITU-T P.911 (“Typical viewing and listening conditions as used in audio-visual quality assessment”). No other light source is admitted, and in particular any light source directed to the screen or creating reflections; ceiling, floor and walls of the laboratory have to be made of non-reflecting material (e.g. carpet or velvet) and should have a colour tuned as close as possible to mid grey.

A.6 Overall test effort and subjects’ involvement
The duration of the test will depend on the number sequences tested in each category / resolution assigned to the test laboratories; in any case each viewing session will not run for more than 20 minutes and the same viewing subject will not participated to the test run for more than six hours in total. The same subject may not be enrolled for two consecutive days. Young humans subjects, equally distributed in gender, are hired, selecting them for an age from 18 to 30 and, highly preferably among University students of scientific faculties. Viewing subjects are compensated for their participation to the testing activities (compensation may be done in money or services).

A.7 Statistical analysis and presentation of the results

The data collected from the score sheets, filled out by the viewing subjects, will be stored in an Excel spread sheet.
Five spread-sheets will be prepared: four containing the results for 4K, 1080p, 720p and 480p (Main profile) and one for 4K (Main 10 profile).
For each coding condition the Mean Opinion Score (MOS) and associated Confidence Interval (CI) values will be given in the spread-sheets.

The MOS and CI values will be used to draw graphs. The Graphs will be drawn grouping the results for each video test sequence. No graph grouping results from different video sequences will be considered.
From the “raw” data subject reliability should be calculated and the method used to assess subject reliability should be reported. Some criteria for subjective reliability are given in [2] and [3].
As an example, the reliability of a subject, could be achieved computing the correlation index between each score provided by a subject to the general MOS value assigned for that test point; in this regard a correlation index equal or superior to 0,75 (computed making the mean of all the correlation values) could be considered as valid for the acceptance of the subject.

References:
[1]          International Telecommunication Union Standardization Sector; Recommendation ITU-T P.910 “Subjective video quality assessment methods for multimedia applications”

Copyright of test sequences

The test sequence and all intellectual property rights therein remain the property of the owner below.  This material can only be used for the purpose of developing HEVC standards.  This material cannot be distributed with charge.  The owner makes no warranties with respect to the material and expressly disclaims any warranties regarding its fitness for any purpose.
Owner of these sequences:
  Owner: Plannet inc.
  Production: Plannet inc. and IMAGICA Corp.

User agrees that the Sequences and all intellectual property rights therein remain the property of the 4EVER consortium members or their licensors.
These Sequences can only be used for internal research and test work dealing with Ultra High Definition, including research and test for standardization purposes. Attributing the work to 4EVER consortium will be done by attaching the following attribution notice to the Sequences:

“Copyright © 2012-2013, all rights reserved to the 4EVER participants and their licensors. 4EVER consortium: Orange, Technicolor, ATEME, France Télévisions, INSA-IETR, Globecast, TeamCast, Telecom ParisTech, HighlandsTechnologies Solutions, www.4ever-project.com, contact: maryline.clare@orange.com. The 4EVER research Project is coordinated by Orange and has received funding from the French State (FUI/Oseo) and French local Authorities (Région Bretagne) associated to the European funds FEDER.”
Your attribution must not be in any way that suggests that 4EVER endorses you or your use of the video sequences.

Subject to compliance with the terms and conditions set forth in the present authorization of use, Technicolor hereby grants to any member of the HEVC and SHVC standardization group (“the User”), a personal, non- transferrable, non-sub-licensable, worldwide, royalty free license under Technicolor owned or controlled copyrights to display (and to copy and modify as technically necessary) the Content solely for the purpose of User’s internal processing, testing and assessment of the Content (or if relevant for the purpose of joint processing, testing and assessment of the Content with another “User”) in order to:
·         evaluate the User’s contributions to the HEVC and SHVC standards (and if relevant to any multi- standard performed in JCT-VC context)
·         evaluate Technicolor’s contributions to the HEVC and SHVC standards (and if relevant to any multi- standard performed in JCT-VC context)
·         evaluate other third party HEVC and SHVC standards contributors’ contributions to HEVC and SHVC standards (and if relevant to any multi-standard performed in JCT-VC context)


The video sequences provided above and all intellectual property rights therein remain the property of the BBC. The BBC is making available the video sequences for use under the Creative Commons Attribution-NonCommercial 3.0 licence.

You are free to use, share (to copy, distribute and transmit) or remix (to adapt) the BBC video sequences, provided that:
  • No-commercial- you may not use these video sequences for commercial purposes; and
  • Attribution- you attribute the work to the BBC by indicating that the video sequences (or elements thereof) were produced by the BBC.  Your attribution must not be in any way that suggests that the BBC endorses you or your use of the video sequences.

Standards committees can use CDVL R&D content within subjective tests to validate objective video quality models (e.g., ATIS, VQEG, ITU).

Individuals and organizations extracting sequences from this archive agree that the sequences and all intellectual property rights therein remain the property of Sveriges Television AB (SVT), Sweden. These sequences may only be used for the purpose of developing, testing and presenting technology standards. SVT makes no warranties with respect to the materials and expressly disclaim any warranties regarding their fitness for any purpose.

Vidyo donates the sequences to the public domain (JCTVC-P0042)

The video sequences provided above and all intellectual property rights therein remain the property of the Kamerawerk. The Kamerawerk is making available the video sequences for use under the Creative Commons Attribution-NonCommercial 3.0 licence.